DEV Community

David Aronchick
David Aronchick

Posted on Originally published at distributedthoughts.org

The Shortest Stave

A wooden barrel built from unequal staves (I love this term, by the way; I've never heard it!) holds water only as high as its shortest stave. Chemists call this the "limiting reagent," but I'm particularly fond of talking about these kinds of constraints in the living world, where things grow ... organically. Long before anyone had heard of a GPU, Carl Sprengel sketched the idea in 1828, and Justus von Liebig made it famous a couple of decades later, focusing specifically on how a plant's growth is set by whichever nutrient is scarcest. Pour on all the nitrogen you want, but if the soil is short on phosphorus, the plant stops growing at the phosphorus line, and the nitrogen sits there, unused, a resource with nothing left to do. So, back to our barrels: you can spend a fortune making the tall staves taller, but the barrel doesn't care. Every dollar spent elsewhere runs over that one short piece of wood and out onto the ground.

During 2026, the price of a stick of server RAM stopped being "Oh, I guess we'll just have to add a little bit to the price of this server, but it's diminimus compared to this Rolls Royce GPU we have in here." Conventional DRAM contract prices rose 93 to 98 percent in the first quarter alone, and TrendForce expected another 58 to 63 percent in the second, roughly tripling in six months. The reason? High-bandwidth memory will take about 22 percent of the world's DRAM wafer starts by the end of 2026, up from 18 percent a year earlier, and turn them into only about 9 percent of the bits, because a bit of HBM eats three to four times the wafer of an equivalent bit of DDR5. SK Hynix, Samsung, and Micron, the only three companies on Earth that make HBM at volume, reported their 2026 HBM capacity sold out under existing contracts, and Digitimes reported in August that most of 2027 is already allocated too. Try to buy a stick of RAM for a gaming PC right now, and you're competing with Nvidia's supply chain. Nvidia is winning.

Everybody saw this coming from the wrong direction. For two years the story was GPUs: how many Nvidia could ship, how long the allocation list revolved around, and whether the shortage had somehow become a permanent feature of the industry rather than a temporary one. Then the story became power. Gigawatts, interconnection queues, and data centers in Virginia are waiting up to seven years for a grid connection. Nobody was watching the beige stick of memory sitting next to the accelerator on the board, because memory has been the least glamorous component in a computer since the 1960s, and unglamorous things don't get congressional hearings. Now it's the thing capping how much of this you can actually build, and a fab that makes memory cannot be told to make more of it by Thursday. Wafer capacity is a multi-year build; wanting more sooner doesn't make it sooner.

The AI industry has been building staves. It built an enormous GPU stack, then an enormous power stack in roughly that order, because those constraints tailed first and made the loudest noise going down. Memory didn't make a loud noise, and so it quietly became the shortest piece of wood in the barrel. The irony is that it's not even HBM's fault! It's just that HBM and consumer memory come out of the same fabs, competing for the same wafer starts, so every additional stack of HBM that Nvidia's next accelerator needs is DRAM capacity that doesn't go into a laptop, a phone, or, this being the actual news story of the last few months, a stick of RAM for the machine on your desk. The industry solved for the two constraints everyone could see coming, and the barrel filled to a different line anyway. One nobody had thought to check.

Even worse, Samsung, SK Hynix, and Micron can't ship a patch. Building new wafer capacity takes years, and the contracts locking up existing capacity run multiple years past that. Which means the memory constraint, unlike the GPU shortage and unlike the power shortage, won't get quietly fixed by a new chip generation or another grid order. It's poured into the concrete of fabs that haven't broken ground yet. The people planning AI infrastructure right now are, whether they've noticed or not, planning against a stave whose length was mostly fixed a few years ago and cannot be argued with, subsidized, or ordered into growing faster by anyone in Washington.

The GPU story and the power story were both real. They ARE both real. Solving them was and is necessary work. Solving the constraint just moves the waterline to the next-shortest stave, and that one is rarely the one the headlines prepared you for. Liebig figured that out, staring at dirt in the 1840s. The AI industry is finding it out staring at a spec sheet in 2026 and paying triple for the privilege.


Want to learn how intelligent data pipelines can reduce your AI costs? Check out Expanso. Or don't. Who am I to tell you what to do.

NOTE: I'm currently writing a book based on what I have seen about the real-world challenges of data preparation for machine learning, focusing on operational, compliance, and cost. I'd love to hear your thoughts!


Originally published at The Shortest Stave.

Top comments (0)