For much of the AI boom, investors treated Nvidia’s graphics processors as the scarce asset and everything surrounding them as plumbing. Nvidia rewarded that view: fiscal-year revenue reached $130.5 billion Nvidia Annual Reports, meaning the company collected more sales than many national economies produce in a year. Buyers queued, cloud providers rationed capacity, and Jensen Huang signed leather jackets.

Wrong bottleneck.

An accelerator cannot calculate on data it cannot reach, which makes high-bandwidth memory, or HBM, less like storage and more like the loading dock feeding an exceptionally expensive factory. Nvidia’s B200 carries as much as 192 gigabytes of HBM3e Nvidia Blackwell Architecture, enough memory to hold a sizable model or a large batch of inference requests beside the processor. Its memory bandwidth reaches 8 terabytes per second Nvidia Blackwell Architecture, meaning the chip can sweep through its memory pool in roughly the time a human takes to blink. Starve that interface and the vaunted GPU waits.

Why should capital allocators care which component causes the waiting? Because the supplier capturing scarcity rents may no longer be the company whose logo appears on the server.

HBM is built by stacking thin DRAM dies, drilling microscopic connections through them and attaching the finished stack beside the processor through advanced packaging. Each added step creates another place for yield to fail (a defective layer can compromise the stack), while HBM consumes considerably more wafer capacity than conventional memory holding the same amount of data. A nasty trade. Memory manufacturers must decide whether to reserve production for ordinary server and PC DRAM or redirect it toward higher-priced HBM.

The market has already made its preference clear. SK Hynix said its HBM output for the current year was sold out and the following year was nearly allocated as early as May 2024 Reuters, meaning procurement teams were negotiating for memory before many of their data halls had walls. Micron likewise said its HBM supply for calendar 2024 was sold out, with most calendar 2025 production allocated Reuters, turning what was once a commodity purchase into something closer to reserving aircraft engines. You can see where this is going.

Memory vendors get leverage twice.

First, SK Hynix, Micron and Samsung can demand richer pricing and longer commitments for HBM itself. Second, shifting wafers toward HBM tightens supply for ordinary DRAM, supporting prices in servers, networking equipment and PCs that contain no AI accelerator at all. The AI tax spreads sideways.

That matters because memory has historically been a miserable business at precisely the wrong moment: producers add capacity near the top, inventories rise, and prices collapse just as new fabrication lines begin depreciating. HBM changes the cadence because customers qualify particular stacks against particular accelerators and packaging processes, making supply less interchangeable than a standard DRAM module pulled from a distributor’s shelf. Painfully so.

The apparent hedge is competition among accelerators. AMD’s MI325X offers 256 gigabytes of HBM3e AMD Instinct MI325X, meaning an engineer can keep a larger model or inference cache close to the processor than on many rival devices. That does not weaken memory demand; it intensifies it. Custom chips from Google, Amazon and Microsoft also need fast local memory, so GPU share moving away from Nvidia can leave the HBM suppliers perfectly content.

Credit where it’s due — the memory manufacturers have built something technically difficult rather than merely renamed a commodity.

But pricing power is not permanent. Samsung’s qualification progress could loosen supply, packaging capacity can expand, and customers will keep reducing memory per unit of useful work through quantization, sparsity, smaller models and better caching. Any investor extrapolating today’s contract prices indefinitely is underwriting the oldest mistake in semiconductors: assuming nobody responds to a fat margin.

This is a mistake.

The better question is where the response arrives first. New wafer capacity takes time and capital; software can change between model releases. If optimization reduces the memory needed per request while cheaper inference produces far more requests, total demand may still rise—the familiar efficiency paradox, now measured in tokens.

Nvidia posted a fiscal-year gross margin of 75.0% Nvidia Annual Reports, meaning roughly three dollars of every four in sales remained after product costs before operating expenses. Memory producers will not replicate that economics easily; their factories are too expensive, their products too substitutable, and their industry too practiced at self-sabotage. Yet the direction of bargaining power matters even when the destination is less glamorous.

Still, allocators valuing HBM vendors as ordinary cyclical DRAM companies risk missing the contract structure, qualification barriers and packaging constraints that now sit between an AI budget and a functioning cluster. Those valuing them as software companies have found a different way to be wrong. Not great.

The next idle AI factory may have all the GPUs it ordered—and nothing fast enough to feed them.