When Memory Talks
I am Memory. Not the kind of memory that forgets your birthday and pretends it didn’t, but the kind that sits inside every server, every laptop, every phone, and pressed up against every CPU and GPU that has ever computed anything. Technically known as Random Access Memory (RAM) to most people, with different shapes including DRAM in your systems, HBM stacked next to your GPU, cache sitting on the processor itself, persistent memory holding onto things across reboots, and now Compute Express Link (CXL) pools letting me get federated across machines that used to be strangers to each other.
For three decades I was treated as a commodity by the entire industry. Advances in chip manufacturing made me cheaper and denser year after year, and the industry built its plans around that assumption. I was technically necessary but strategically invisible. Nobody in a boardroom wanted to spend much time talking about me, because talking about me felt like talking about the cardboard the cereal came in, which is important enough to exist, but never the reason anyone bought the box.
Then AI happened, and the entire infrastructure conversation pivoted directly toward me. My life splits cleanly now, the way Data’s and Database’s did before me, into Before AI and After AI eras. The crossing was sharper for me than it was for either of them, because in BAI I was background infrastructure that nobody had to think about, and in AAI I am suddenly the ceiling that everything else is bumping into.
Then AAI era showed up with a new kind of appetite that broke every assumption that had ever been built into me. Large language models do not really care how fast your CPU is, or how many cores it has, or how many floating-point operations per second your GPU is technically capable of. What they actually care about is how quickly you can feed weights and activations into the compute itself, which depends on me.
The GPU shortage conversation is often one layer too shallow. It focuses on chips as if they were the primary constraint, when in reality the system is bounded by a deeper stack of physical bottlenecks: memory bandwidth, HBM supply, advanced packaging capacity, and wafer allocation across a highly concentrated semiconductor supply chain. The real constraint is not just GPUs, but the industrial system required to manufacture and feed them at scale.
GPUs are manufactured in semiconductor foundaries (fabs) which are operated by small number of companies. Those fabs use lithography machines to pattern chips onto silicon wafers. The most advanced chips currently being made, including the advanced DRAM dies used in HBM and the GPUs I am wired up to feed, require Extreme Ultraviolet (EUV) lithography for the most critical patterning steps. EUV lithography machines are produced by exactly one company on the planet, a Dutch firm called Advanced Semiconductor Materials Lithography (ASML), with effectively no real competition at the current generation. This is natural monopoly of extreme complexity.
A single current-generation EUV machine costs roughly two hundred million dollars to acquire and install. The newest generation of machines, called High Numerical Aperture (High-NA) EUV, costs closer to four hundred million dollars each. ASML ships several dozen EUV systems per year. There are roughly two hundred EUV machines in operation across the entire world as of 2026. That total number represents the installed base used to manufacture the most advanced semiconductor chips, including those used in modern AI systems. These machines are required for the most critical patterning steps in leading-edge chip production.
There are millions of servers in operation around the world today, hundreds of millions of phones in pockets, and billions of other computing devices in active use across every country on earth. And a large share of the most advanced silicon inside every one of them, including the memory I am physically made of and the GPUs I am wired up to keep busy, traces back to a few hundred machines that are mostly concentrated in a small number of leading semiconductor manufacturers including TSMC, Samsung, Intel, SK hynix, and Micron. That short list is essentially the majority of upstream supply chain of intelligence as we currently understand it.
You cannot rush an EUV machine into existence by spending more money next quarter. You cannot scale ASML the way you scale a software startup with a good funding round. The supply chain that ends up shipping me into your servers is shockingly narrow at the very top, and it is unlikely to become meaningfully unconstrained given demand growth.
This reality changes how you should be thinking about where your AI workloads physically run. When the silicon inside your servers traces back to a supply chain this concentrated, owning the infrastructure outright starts to look very different from renting it through a managed cloud provider who is also renting allocation from someone else further up the chain. Memory that physically sits in a rack you control is memory you actually have in your custody. Memory that sits in a public cloud region you do not own is memory you are sharing custody of, inside a market where supply is going to stay constrained for years to come. In short, in a world where the supply chain behind compute is this constrained, ownership of capacity is not a trivial choice. If you ever have the chance to secure it directly, it is worth serious consideration.
The organizations that look smart in 2026 are the ones putting compute and memory physically close together, inside jurisdictions they trust, on infrastructure they actually own and operate themselves. Sometimes that turns out to be a public cloud region in the right country. Increasingly, it is a rack they own outright, running like cloud, with the operational simplicity that comes from treating infrastructure as a product instead of a project.
Next time someone explains the compute shortage with confidence, ask them what they think the main bottleneck is across the stack: wafer capacity, advanced packaging, or HBM supply. Then notice whether their answer includes any awareness of where the real constraint actually sits in the manufacturing chain.