What the freaking hell.
GPU is mostly fine and can probably reuse that but it’s Vega so not the greatest I was due for an upgrade but the prices of everything is ridiculous.
The thought of pc-of-Theseus’ing my busted ass PC is just sad.
Thank god I bought a PS5 Pro before Sony started jacking up the prices.
I remember when buying a $300 GPU was extravagant and gave you state of the art graphics.
Pick any B550 with a VRM heatsink for $90, a 5600 for $190, and you are back in business.
Get a 5070 or 9070 before GDDR supply is completely gone if you want an actual upgrade.
As for GPU price, that Vega 64 was $500 in 2018. Even a midrange GPU hasn't been $300 since the Radeon RX 480 launched a decade ago.
The PC I replaced when I bought this one was such a huge improvement.
My options basically boils down to fixing trash that’s falling apart or a sidegrade at best for the same price as what I bought a beast of a PC for 10 freaking years ago. Depressing.
Point is, quite a lot of hardware can still do a lot. I have a Ryzen 5500U-based laptop that still outperforms the PCs of most people that I know (non-tech people).
Though in my case I feel like I have to get a loan to buy a workstation + an actual gaming PC for my wife. That all is beginning to look like at least $15k... yikes. But I am not sure I can wait until 2028 to see if things _start_ normalizing... Depressing indeed.
surely they cant doing the same for 2028 right ??
My solution is to see what four or five-year-old equipment I can buy that will let me run local LLMs. I may only get six or seven tokens per second out of an i7, but it's a start. And best of all, I can turn the machine off when I'm not using it.
IMO, Migrating to small-scale local LLMs would be a significant improvement over using data centers.
LLM serving is most efficient when you batch a lot of parallel requests together. Data center solutions also have the advantage of collecting queries from around the globe, so the hardware can be utilized around the clock.
Having everyone serve their own local LLMs would produce a lot more memory demand. Not less. The same memory would be idle most of the time, and when it was used it would be used for 1 person instead of a batch of requests.
There are other reasons to run local LLMs, but solving hardware demand problems is not one of them.
This shift is probably inevitable, but it will vary significantly by region depending on prices for electricity. Look at, for example, the difference in fundamental homelab build recommendations between Germans and just about anyone else. Electricity prices in Germany are so high that even a now expensive Raspberry Pi or other ARM board is often preferred over Intel/AMD builds due to low power draw (especially low idle power draw), an effect that adds up for a machine running all the time over years.
With local LLMs and the GPUs to run it, especially if you want a model available to you all the time and can remote into your local network to use it whenever you want, there's no escaping much higher power draws, even at idle. Wherever electricity is expensive, the electric bill can be a prohibitive barrier.
Eg: I shaved ~40W off the idle load on a server (250->210W) by doung nothing more than removing the redundant supply
A college buddy used to work at Motorola (I'm naming the company because they wouldn't mind this story being shared) back in the late 90's or early 2000's. They had redundant power to their campus, bought from two different companies, coming in on opposite sides of the campus, so that even if some backhoe operator cut a ground-based power line somewhere, they wouldn't lose power.
And yet, one morning, the power went off all across their campus. After a little investigation, they sent pretty much all their employees home at noon and told them "take the afternoon off, don't come back until tomorrow, you wouldn't be able to do any work anyway". Turns out that although the power lines came in at opposite sides of their campus, somewhere a few miles away both of the power lines feeding their campus ended up running through the same underground conduit. And yes, a backhoe had managed to cut that conduit and break both of the lines they depended on at the same time. They had a single, VERY non-obvious, point of failure, and the backhoe had unerringly homed in on that SPoF.
I'll second this. The combination of Whisper + LLM makes speech recognition fantastic. I occasionally have arm pain from typing, and this is a Godsend.
I don't use it to write code - but in my experience stuff like emails + docs was the greater source of pain (one generally types slower while coding).
The average consumer isn't purchasing a lot of products with a lot of RAM every year.
Their 8GB of RAM phones will go up in price a little bit, but people aren't buying phones every year or even every other year.
So if the price of the 8GB of LPDDR went from $40 to $160 and it's all passed on to the consumer buying a new phone every 4 years, that's an extra $30/year in spending.
Most adults I know now use their phone for everything, so they're not buying new computers and laptops. They can wait 3 years for the market to settle before upgrading those, too.
Even I'm a heavy buyer of tech products and RAM, and I would bet you that my family's annual food bill fluctuates by more than what I've had to pay for RAM prices growing.
Many smaller companies will have issues getting enough, but even they could buy something with memory and rip it out. When you're willing to pay a hefty multiple, you can outbid other people. The shortages are not nearly bad enough to escape standard supply and demand.
>Most adults I know now use their phone for everything
You mean the cloud for everything. A lot of phone tasks share their computational and storage workloads off in the invisible ether that has actual computers with real costs behind them. Amazon has no problem with bumping up their costs in order to pay for their fleet of servers. I mean, what business doesn't use the cloud these days.
In the shorter term, you’ll probably see more “base” models on offer with some of the most intensive features disabled. There’s still going to be all of the telemetry stuff that makes them money through third parties though.
Isn’t it just… prices?
Total imports are maybe 14%? But not as highly tariffed.
Your computer already supports that. It's called "swap".
CPUs have L1, L2, L3+ caches too after all.
The M10 are over produced and only PCI-E 3.0 2x lane devices but 16GB of DDR5 is going to cost over $200.
You can do the same with USB-C M.2 adapters but the adapters cost more and eww USB for Swap, hope the device never drops.
https://www.amazon.com/GLOTRENDS-Adapter-Aluminum-Heatsink-P...
I suggest using an 1x PCI-e adapter because putting a M10 in a M.2 slot is likely going to only have 2 of 4 PCI-e lanes in use. You have to figure out the tradeoff of using a Swap on your, now very expensive as well, Flash storage that may have flashier sequential read and write speeds but doesn't have the ability to sustain random read and write performance which is what a Swap is going to use the most.
There are faster Optane drives that use 4 PCI-e lanes or even PCI-e 4.0, but they get pricey quickly, so you should just buy ram then. The deal is the 16GB Optane on an unused 1x PCI-e slot.
In any case DDR3 isn’t nearly as bad price-wise and would have much better energy efficiency than stacking 1–2 gig sticks.
For starters, The slowest sticks of DDR(x) are often slower than DDR(x-1). The issue is never capacity, but rather, performance.
The real nonsense? USB? It's a mess. Pick a USB cable and buy it from ANY retailer, let's make it easier, buy a USB-C cable. What is the data rate (depends on cable quality and length), Does it support power delivery? If so, how many watts? (~5W requires a very different cable from ~230W), how do you know from simply looking at the cable? If you buy a cable, how can you tell what it supports by simply looking at the connector? Imagine having a box full of USB-C cables. Could you tell me how fast each of those cables are? (The cables themselves don't! Many don't have any markngs at all, and if they do, the markings could be fraudulent)
RAM does not have that issue. A stick fits or it doesn't. Sure, there are a huge range of speeds, however, that range has a ballpark (JEDEC defines the ballpark, the "cartels" make the memory and push out some faster stuff).
While there are definite exceptions (I was bitten by one recently), you can generally plug in a DDR5 DIMM and expect it to work in the system.
The same cannot be said for USB. Some USB cables ONLY deliver power. Some only work with certain devices. I've USB-C (!!!) cables that only work with the devices they are shipped with, and even more annoyingly, the both may be true! ASUS (!!!) ships MOTHERBOARDS that can only pair with vital hardware via very specific USB-C cables (Strix Hive) and "GLORIOUS" ain't so glorious. Their mouses warn you not to use other USB-C cables, and they are right, depending on the make/model/generation, you can brick your "glorious" peripheral.
No, USB nonsense needs to stay far away from anything else.
Also, the reason this is a huge issue is because the memory makers are cartels. Only a few of them exist, they gang up and bully EVERYONE and nobody has ever invested money to create a competitor due to this, except China, which of course means that the U.S. and portions of the E.U. are insta-banning/trying to insta-ban, even though it is really freaking hard to install spyware on a memory module.
Okay, let's assume it's a USB-C cable and not just something USB-C shaped that pretends to be one.
> Does it support power delivery?
Yes.
> If so, how many watts?
Always at least 60W (3A). Up to 240W if it's e-marked.
> What is the data rate
Depends whether it's a USB 2.0-only cable (480Mbps) or a complete cable (20Gbps). Possibly more in Thunderbolt or USB4 modes - cables that handle those will be marked (both visually and with an e-mark). In any case, pretty easy to check just by plugging things in.
If we're counting fake products then we need to count those RGB sticks that have no RAM in them.
Most recent one I bought a multi-tb hard disk, but got an old multi-gb disk shoehorned into legitimate box. Couldn't just get a replacement - had to buy one for more money since that drive's price had increased.
HP, Acer and Asus are now using them
https://asia.nikkei.com/business/china-tech/hp-asus-and-acer...
Pretty sure those devs were saying the same thing way before 2017 as well, which seems to be ~ the last time RAM was this expensive.
Now RAM in the cloud, now you're really paying the Java premium.
Since that peak in 2014 until today, the prices were lower than now. The peak in 2017 was more than 10% lower than today.
To reach permanently higher prices than today, we must go backwards until 2011.
So we have already regressed at least 12 years into the past, but more likely 15 years, and it is unknown how much more we will regress.
1. Make extra greasy pizza. Nauseating-level of grease.
2. Put pizza through De-Greasinator 5000.
Sounds efficient.
Were devs supposed to optimize RAM for a shortage that might come? I’m guessing you had enough foresight to stockpile RAM?
1) .net app, one text field, one button; private bytes 22mb, working set 27mb
2) native app, two text fields, three buttons; private bytes 1.2mb, working set 7.3mb
Posting instead of researching in hopes someone smarter can chime in, because I'm lazy.
Also, you have Meta and Google investing in simplified performant versions of their stack for developing countries which is similar.
What's funny though, I see people that sometimes says it's cheaper to build the app they need in a single prompt than to search for it.
Moore's law is alive and well.
Or lease a laptop from Apple if faceputer's aren't as much your style.
"The original retail price of the computer with 4 KiB of RAM was US$1,298 (equivalent to $6,900 in 2025)[21] and with the maximum 48 KiB of RAM, it was US$2,638 (equivalent to $14,020 in 2025)"
For the skeptical ones: Somewhere people are already being cut off from state and financial institutions without proprietary software with bundled security certificates, applications are bound to collect data on the environment they are used in (other apps) and organizations deliberately limit access to their services or cripple them without their applications (can't do things from a generic web-browser).
The words after this point make you less convincing, not more, as they only serve to confirm the egregious slippery slope of your reasoning.
(PS: It is time to let ASML sell to CXMT.)
Other manufactures like Micron do have aggressive expansion plans, but fabs take a long time to build and bring online - multiple years.
https://www.guru3d.com/story/cxmt-dram-capacity-reportedly-s...
Traded as CUSU for end-users: https://en.twsc.com.cn/enCUSU/index.html
Why? High demand and low supply is very good for shareholders and make indexes go up, their goal is not to meet demand, they want exponential qurartly revunes
To me it seems like they’re just riding the AI hype out, expecting a crash after, and not willing to risk a lot of capital on someone else’s bubble.
If memory is now upstream of so many companies, then maybe a higher regulatory involvement would be appropriate (like power companies).
Specifically, and it could fix all of this, capping the max % of their supply they're allowed to sell into the AI space.
The arguments for infrastructure being managed more tightly by the government is because it's a natural monopoly that everything depends on. Memory isn't a natural monopoly, if we found ourselves in a spot where we all depend on 3 companies then that's not a law of nature, we need more companies.
The big three are also expanding their capacity. IMO, it seems to make more sense to sell equipments to those who could make better, efficient use of them. ASML has a backlog of €38.8 billion and is fully booked for 2027.
I had previously ordered from Aeonfly, but they cancelled my order--likely because they wanted to raise the price.
Do deliver and they don't get paid because the buyer purchased with money that doesn't exist.
Lose-lose situation!
What we got in 2008 was central banks (eg Fed and ECB) willfully collapsing nominal GDP in their economies. Have a look at the dot-com bust or Black Monday for comparison.
What on earth are you talking about?
The credit risk was a very real problem; Kaupthing, Anglo Irish, RBS, Lehman etc.
In 2008 they didn't. Instead they actively tightened monetary policy by eg introducing interest on excess reserves. The ECB even hiked interest rates.
Companies defaulting on debt doesn't need to bring down the economy.
For the US, you can also see how the construction sector had been winding down for years (eg as measured in construction employment) without an impact on overall unemployment. The crisis was entirely avoidable.
See https://www.cato-unbound.org/2009/09/14/scott-sumner/real-pr... for a bit more background.
Take what's currently happening to Flock cameras as an object lesson. People are so fed up with Flock that they're cutting their cameras down en masse, sometimes wiping them clean from entire jurisdictions. Tens of thousands of other people in those jurisdictions are cheering them on, even offering alibis for them before they've been caught. Law enforcement doesn't seem too keen to catch them either. You absolutely can be so hated that the law will not be enforced to protect your property.
If AI companies don't slow down and build some good will, AI data centres are probably at risk. They'd probably have been targeted before Flock were they not harder targets. A higher degree of cooperation and organization will be required to sabotage them, but it would be foolish to believe it won't eventually happen.
Monopolizing memory production for another year suggests that lessons have not been learned, the mad build-out will continue, and we're headed for some truly crazy stuff.
And if AI is putting people out of jobs and making everything expensive - or that's just the perception - that's gonna set AI back for years.
Civil disobedience? In this economy?
It seems inevitable that these lost cameras will simply be upgraded to drone fleets. Not an identical replacement, but still "good enough" for the original purposes. Had they anticipated this problem and skipped the first phase, it might have been harder to manufacture consent for privacy invasion. But now anybody who was okay with the original cameras feels like a victim having them torn down, which makes it likelier they'll be okay with drones everywhere.
> Tens of thousands of other people in those jurisdictions are cheering them on, even offering alibis for them before they've been caught. Law enforcement doesn't seem too keen to catch them either. You absolutely can be so hated that the law will not be enforced to protect your property.
At least one third of the country is still okay with them, because they believe it helps rid them of illegals and deter crime, with downsides small enough to ignore, in their mind.
In sum, we don't have as much control over these things as you seem to believe. Capital was always in control with its crude but effective propaganda pipelines to manufacture enough consent. Now imagine those pipelines becoming even more effective thanks to AI. Not looking good.
Who cares about those people? I would argue even they themselves hardly do. And sure, they'll be even more okay with even worse stuff, and that even worse stuff will be even more hated by even more people, whatever the people who are "okay with" (that can mean ignorance, obedience, or evil intent, but it's not justification). If they can be okay with mass surveillance they can be okay with other things, so they're basically a wash, they're just okay with stuff.
> At least one third of the country is still okay with them
Then they can be okay with the majority saying you know what, we're sick of you, you can either stop having your toys or you will be segregated from us and live under your own surveillance. "Being okay with something" isn't some magical wand or a fortress, it's just a limp shrug. It means "pass".
> In sum, we don't have as much control over these things as you seem to believe.
I don't see the connection with the sentence before that, which is a minority being "okay with" something and a majority willing to fight and out there doing it. It's not about controlling what a minority of people are okay with, it's about changing what is done. The people who are okay with stuff can go read a comic book until it's over, how does that not solve anything you raised about them? Or are you saying their being okay with stuff should be respected or even heeded? Couldn't they just make that claim themselves, if they wanted to?
Imagine a raging house fire, people stumbling over each other trying to help, and some guy steps up and goes "WAIT!" they all look at him, sure he has a major contribution for interrupting something like this, a good idea, a plan perhaps; and he goes "... I'm actually okay with this."
It's not the majority, though. Roughly speaking, among adults, we have:
* 1/3 who are willing to vote for the red flavor of lesser-two-evils (some of them don't even register it as evil)
* 1/3 who are willing to vote for the blue flavor of lesser-two-evils (some of them don't even register it as evil, including when it literally manufactures consent for the first third to win next time around)
* 1/3 who see the above as a uniparty of capital interests offering the illusion of democracy through "close" elections as the bullshit that it is.
I wouldn't do that unless you want your airspace filled with birdshot
Nobody wants this robber baron shit
It's all fair game IMHO, except the elephant in the room which is that the AI company bubble might pop.
It is also a term use in general trading anywhere traders open and close their positions quickly to make profit off an upswing which might be in part caused by traders doing this. If done with inside knowledge it is illegal, in this case more commonly referred to as “front running” (which is what many suspect happened with numerous ahemfortuitous/serendipitous position changes around announcements of changes in the state of the US vs Iran situation).
> Also, these companies supposedly intend to actually use the RAM they ordered, not reselling at some inflated price.
As well as the companies ordering for the purpose of DC roll-outs and upgrades, there will definitely be some buying purely to sell at a higher price a short time later.
> It's all fair game IMHO, except the elephant in the room which is that the AI company bubble might pop.
[and with reference to the sibling reply to this: “Then we'll have a lot of cheap RAM.”]
I don't think it will pop with a sudden bang, and even if it is it won't affect these memory sales or probably those for 2028. The first real effect will be those using the AI tools suddenly finding themselves needing to pay a lot more for them as companies offering those services deal with reduced access to “new” credit from the circling investment pool shore up their fanciful accounting that way instead. The DC builds/upgrades that are currently planned will happen but more funded by end users and less by the investment cycle. New plans might not be actioned at the same scale, but it will take a while for that to affect component prices as there is a lot of lead time involved.
A year ago the SOTA was GPT-5/Opus 4.1/Gemini 2.5 Pro, two years ago Sonnet 3.5.
Looking at Opus 5 for SOTA performance and GPT 5.6 Luna for a viable cheap alternative, AI is much more capable now.
Honorable mention: GPT 5.6 Sol on Cerebras, capacity limited to a few customers, is supposedly serving 750 tok/s.
Compared to 90 tok/s for non fast mode 5.6 Sol, 56 tok/s for Opus 5, and 190 tok/s for 5.6 Luna.
I am very curious about the next generation of models, GPT-6/Astra is rumored to launch still in August. Not sure what is the state of Anthropic's next Fable checkpoint.
If these models also deliver significant improvements, I really do not see how one could seriously still argue among the lines of AI being a scam, and the demand not being there to support the size of the investments.
I don’t think most people are saying it is purely a scam (well, some do but I don’t think they are to be taken seriously). What all these talks about circular financing and VC money are saying is that demand cannot sustain the sector long-term, not that there is no demand. People are certainly happy to pay say $20/month for whatever AI chatbot, but would they still pay if it were $200/month, which is closer to the actual costs.
There are several other details that point towards a unsustainable projections:
- measurable benefits from AI-ifying companies are nowhere near what is commonly believed. AI providers are hoping that they can keep the show going until the models are good enough, essentially faking it until they’ve made it, but that is not a given. It is also unstable because it is susceptible to change in public opinion.
- permits for new datacenters are not going to become easier to get as public opinion keeps turning against them. They will have to concentrate in friendly regions, which will add cost (more demand for the same location, plus interconnection for network and electricity, both of which can easily become bottlenecks).
- the electric grids are not ready for all those planned datacenters, so something will have to give. Building more and more on-site diesel generator in times where oil supply is so constrained and random is not very sustainable either.
- eventually the loans will come due and if earnings do not match there will be a, possibly severe, correction.
If we look at the historical example of the dot-com bubble, the crash was not caused by no demand. It was just caused by over-estimated demand and too much money going to a single sector of the economy. We still use the Internet, and it is still hugely important, but the correction was still severe and real people lost real money.
- Chatbot usage in $20 plans is likely nowhere close to a $200 cost.
Inference cost has come down rapidly, with reports from July claiming OpenAI can now serve all of the logged out ChatGPT traffic on just a few hundred GPUs.
- Measurable benefit: I doubt anything of value is being measured. MS Copilot with GPT 5.5 Instant processing SharePoint files? Developers using AI as a fancy autocomplete under a "I review every line" regime?
I rely on my personal value judgement, based on 30h/week I spend using AI outside of my regular job.
I am developing a mobile app, competing with companies with millions in revenue and entire dev teams. I know it is viable. Others lag in effective adoption, their opinion is likely to change soon with even more capable models.
- data center projects in the US: They look to me to mostly be constructed in the most remote backwater. If even there projects with such moderate environmental impact cannot be realized, that would be an embarrassing policy failure
- the grid: I think that one is true. AFAIK the constraint would be gas turbines, not diesel, and oil supply is not structurally constrained
- the loans: Anthropic is rumored to have become profitable earlier this year because of large growth in enterprise revenue. It does not look so bad to me
To add to that, as an ex-software dev that mostly works in a semi-unrelated field now (and can only code as a small part of my job): I think a lot of software devs are underestimating what someone with reasonable technical skills and specialised domain knowledge is able to create with AI.
I honestly wouldn't be surprised if, in 5 years' time, the majority of software used by (e.g.) potato farmers was primarily created by other potato farmers. The code might still be less elegant but I think it will be easier for the potato farmer to iterate with an AI than outsourcing to a dev firm.
I have little doubt that a competent potato farmer could possibly implement and sell such software today using Opus 5, Fable, or 5.6 Sol.
Now what I am going to say next may sound a bit crazy, and it is outside of my area of expertise. I hear the recent mathematical breakthroughs made by GPT-6/Astra are field medal worthy discoveries.
What if the current trajectory of improvement holds for another year or two?
Is it impossible that LLMs gain superhuman ability to reason over a large number of constraints so that they can make novel breakthroughs unimaginable to us today?
What I hope to see in five years is not potato farmers writing software, but programmable immune cells that safely kill cancer.
Perhaps I am an optimist.
AI grows and grows and becomes god and kills us all.
How did it all start? Oh we invested all our money in AI.
We rewrote all our Python code in Go because the RAMs too damn high!
Nothing prevents wholesalers from selling their memory or memory futures at a loss, as that is a risk that is justified by equal possibiltiy of the memory appreciating in excess of inflation.
Besides the AI craze, to me it sounds kinda reasonable. Most large buyers (apple, public clouds, OEMs) do forecasting and orders in advance, right?
Not sure if that's what you've meant...
I would expect that old DDR4 machinery to end up somewhere else, still chugging along. Maybe someone buys the old production and keeps chugging low quality DDR4 or something.
To me, it sounds like that didn't happened. Fabs that were producing legacy DDR4 are still producing it, just selling it higher and surfing on the consumer price increase of DDR5. As DDR5 became more expensive, people looked to build old AM4 PCs with the older memory, and the market just raised the price so that option would also cost more. IMHO DDR4 price increase has nothing to do with production.
I built a new gaming PC with DDR4 two months ago and it's pretty great. For servers RAM speed usually matters even less.
There shouldn't be much issue converting those 16nm DDR4 fabs to DDR5, it's really just an issue of swapping out the masks. That's assuming they haven't all already been converted years ago.
Even when there isn't an overlap, moving to the next process node usually isn't so much about replacing machines, but adding more machines. To oversimplify, you might not be able to convert one older DDR4 fab to DDR5, but sometimes you might be able to convert two older DDR4 fabs to one DDR5 fab.
Such conversions might not be the most cost-effective option in the long run (that unmodified fab could have kept selling DDR4 for years), but conversions might be significantly faster than building new fabs.
Like everything related to it being a clean room; All the logistics robots which move wafers around; All the machines that deposit layers of material like CVD/AVD/Sputtering; The various machines that do the actual etching; The furnaces; The inspection and quality control equipment;
Also, absolutely everything to do with dicing and packaging. But that's usually already at another factory (often a completely different company) due to how process agnostic it is. Same applies to growing and prepping the silicon wafer.
Game: Just buy up all the RAM...
> As discussed previously, the ramp of HBM production will constrain industry supply growth in non-HBM products. Industrywide, HBM3E consumes approximately three times the wafer supply as D5 to produce a given number of bits in the same technology node. With increased performance and packaging complexity, across the industry, we expect this trade ratio for HBM4 to be even higher than the trade ratio for HBM3E. We anticipate strong HBM demand due to AI, combined with increasing silicon intensity of the HBM roadmap, to contribute to tight supply conditions for DRAM across all end markets. As the memory industry is still recovering from the challenging environment in 2023, this tight supply environment will help drive the considerable improvements in profitability and ROI (return on investment) that are needed to enable the investments required to support future growth.
https://investors.micron.com/static-files/4550f98c-1054-4847...
The problem is that they need actual chips tomorrow. I fear we are reaching the point in the semiconductor industry where, in order to sustain the revenue growth propping up their valuations, they're going to have to start selling future chips that cannot possibly be physically produced.
If there is demand for N chips at $X price, but you only have N/2 chips, then half the people aren't going to be able to buy them. The people who really need the chips, and would be willing to pay a lot more than $X for them, will be competing with people who are only willing to pay $X for them. You will end up getting scalping and shortages and hoarding. Since people know that there are enough people willing to pay a higher price, everyone will try to buy them at $X, even if they don't need them at all. Just buy them at $X, and immediately sell it to one of those companies willing to pay a lot more.
The market is extremely inefficient... since the manufacturer isn't charging enough, you get way too many people trying to buy them from the manufacturer.
So you find the price where the demand is N/2 chips, and everyone who is willing to pay that much will get one, and there is no profit for scalpers so the only people who will buy them will be actual companies that need them.
I don’t believe scalpers are an issue in B2B; these aren’t concert tickets being sold to the general public.
It’s clear you grasp the economic theory as it’s taught, but my comment was meant for you to question it. If someone is hungry you can charge more for food. The market allows and encourages it. But don’t mistake that for “willingness” and don’t mistake raising the prices purely to get extra money from the exchange as some inevitable law of the universe. You’re welcome to love the concept, but don’t whitewash it.
Demand is elastic. If the price of potatoes rises to $1000 / kg most people will switch to using pasta or rice for their dinner instead. There might be a shortage of potatoes at $1 / kg, but not at $1000 / kg.
If my 7-year-old computer were to die tomorrow I would've normally replace the entire thing. It has had a good life, the newer generations of hardware are a decent bunch faster, and it would be nice to get some additional features. With current RAM prices? No way I can afford to upgrade to DDR5, I'll have to get a replacement AM4 motherboard / CPU to fix it.
> I don’t believe scalpers are an issue in B2B
They are called "speculators". DRAM is a commodity and is traded no different from, say, potatoes. It's why there are companies like DRAMeXchange. DRAM module assemblers like Kingston, G.Skill, and Corsair will buy chips from whoever gives them the best deal. Similarly, anyone assembling hardware using DRAM will have stockpiled it when the AI boom became noticeable - with some almost certainly selling it off now that it has become too valuable to use for cheaper products.
In a lottery, the goods are misallocated to a bunch of people who aren't getting the most value out of the goods, creating economic inefficiency. A bunch of GPUs would be sitting in warehouses waiting to be resold (either when the price goes up, or in the inevitable black market), or in my basement screwing around with them, instead of being deployed in a way that the most people can benefit from creating maximum economic value from their deployment.
Your gaming PC isn't as economically valuable as JPMorgan using AI for fraud detection, for example.
Moreover, if you are appointed god and force everyone to sell GPUs at one dollar just because you want cheap GPUs, then this is the last batch of GPUs that will ever be produced and you'll have a shortage until the end ot time.
Sometimes this works and sometimes it does not.
If you happily sell apples for $1 and see a hungry person walking towards you, must you raise the price?
I think that basic thought gets lost sometimes when people talk about shortages causing the price to increase. The shortage didn’t cause anything, some executive decided they want more money. That’s all. There’s nothing inherent in the system that requires it. Whether that’s OK or not is up to the reader, I just think people lose sight of the reality and talk about it like it’s gravity, rather than simple decision making.
I’m confused because this reads like a denial of basic economics. If the price is high enough, the chip will be produced for you.
Do you mean because of the production lead time, higher prices won’t result in increased production? Commodities like corn have been managing this for a long time… what’s special about chips?
What is your actual argument?
Samsung, SK Hynix, and Micron have a combined market share of 90% - with most of the rest being a very new-to-the-market CXMT. It is a cutthroat market which behaves like a stereotypical "pork cycle". Semiconductor fabs cost billions to build and take years to complete, so you better be damn sure you have buyers before you start constructing one. You and your competition overestimated the demand? You have to pay back the construction cost, so you're now in a race to the bottom and one of you is going bankrupt.
Ever wondered where Intel came from? They started out as a DRAM manufacturer, which dominated their revenue well after the introduction of their first microprocessors. But in the early 1980s the glut of supply from new Japanese manufacturers made it so unprofitable that they had to ditch the memory market altogether. The stories of Texas Instruments and Motorola aren't much different. And that's not even mentioning the likes of Mostek, which once held a 85% market share and was dead less than 5 years later! Oh, and those Japanese manufacturers? All gone, pivoted like Intel or died like Mostek.
So no, the three remaining DRAM manufacturers aren't going behave like headless chickens and start ordering new fabs just because there's a bubble causing a temporary demand peak. Unless those AI companies are going to pay in advance, in cash, for an entire fab, they'll just have to wait and deal with the price increase.
So it's more "what's special about corn". It is also fairly hilarious to claim the parent is denying basic economics and then bring up corn as an example of having successfully managed economics. If the scales were not being thumbed, and "basic economics" were in play, corn would be in very very bad shape.
In the case of DRAM, there is an incredibly long history of these gloom/glut cycles, and they have stayed roughly the same timeframes (~3 years) since the 1990's.
Almost all the ones who have survived this long are either in the same kind of boat as corn - protected in various forms from the downside - or don't increase production and get caught out until they are absoultely forced.
The very temporarily increased profit is not worth going bankrupt for - they make more money long term by being very cautious and know this.
There are a near infinite number of economic studies you could look at (and several sibling comments cite some) - DRAM manufactuers don't chase the price and probably couldn't anymore if they want to.
None of this denies basic economic theory, of course, since economic theory is not exactly "rigorous", even to the degree it could be (IE even the parts that are pure analysis of data rarely reproduce!).
Memory is just too useful now that you can use it to drive cars and write code.
Which directly leads to the next big development: all the big players are investing in silicon with "baked-in" models, like [0,1]. Turns out you don't need an expensive general-purpose GPU with heaps of RAM to contain a model when you can make a custom ASIC around one specific model! Why spend a fortune on DRAM / HBM when all you need is some finetuning parameters which are easily stored in on-die SRAM?
[0]: https://www.theregister.com/systems/2026/08/06/amd-acquires-...
[1]: https://thenextweb.com/news/google-frozen-chip-gemini-silico...
In any other business choosing (2) would mean someone else swoops in and steals all your business. It doesn't look like this is at all possible for memory fabs.
"If the price is high enough" is of course technically true, but the scale of what high means in this context is important.
> Now, you could raise prices to the point where you destroy demand.
You know supply and demand is like a curve right, you can find an optimal equilibrium? It’s not a cliff that you can fall off.
Maybe you could pay someone else who is less desperate to part with some, but that does not increase supply.
Nine women cannot produce a baby in a month, even though you could average about one per month if you wait about nine months.
That's really interesting, and I wonder why? I believe HBM has redundant ECC bits by default, which would add a few %, but other than that, is it just that the yield is much lower due to die stacking? Of course, this is a 2024 document so things may have changed a bit since.
In order to get high bandwidth, you want memory as close as possible to the GPU. The more trace lane length = signal loss, bandwidth loss.. HBM is compact, and so you can stack 24GB modules, 8 around a GPU die.
If you tried to do that with normal memory, you need like 64 modules. So a a TON of traces more that all need to be equal length, and because so many = far away from the GPU = less bandwidth.
The issue is like stated above, its a process that waste a ton of wafers. Wafers that can make easily 3x more normal memory.
Intel with "Crescent Island" is trying to make a 160GB card using LPDDR5x memory but the bandwidth is only ~700GB/s.
The base interface die has eight or 12 memory dies stacked on top of it. The wafer cost of that interface die is therefore small.
If the process of die thinning and TSV stacking reduced yields by a factor of three, nobody would consider HBM mature enough to put into production, and especially not mature enough to be increasing stack height from one generation to the next.
Trace length has approximately nothing to do with die size. If anything, designing for shorter traces means you can get away with smaller PHYs at either end.
What might go some ways toward explaining such a huge difference in die size is that the TSVs themselves take up significant die area and must be fairly numerous to carry both a large number of signal wires and all the power and ground required by the stack. But it's wildly implausible that the die area consumed by the TSVs would be significantly larger than the die area consumed by the memory arrays themselves, or that anyone would build a memory die where the memory array was not a large majority of the total die area.
If there's any truth to that ~3x higher wafer requirement for the same number of bits as compared to DDR5, it must be a combination of several factors and probably includes something non-obvious and dubious, like counting all the area of the passive interposers that go between HBM stacks and GPUs (those interposers aren't competing for the same fab space).