When Network Talks
Network. You might heard my name in different contexts, because I come in different shapes and forms. Like your network of friends, Wifi network, linkedin network and so on. Network simply means connection between two entities which faciliate the flow of information.
Today, I will talk about my most complex form, which is used by large organizations to move the information across thier infrastrcture. Like humans, I also got impacted by AI in last two years.
Over the last three decase, I was designed around north-south traffic as the primary pattern. Those were mostly happy go lucky days of my admins when users pulled from servers, servers responded, and everyone was home by five. Bandwidth planning was straightforward and mostly linear in nature. I was playing important and crucial job from a long time, but never seen the amount of pressure I had in last two years.
My life splits cleanly into Before AI and After AI eras, similar to my other infrastrcuture peers including database, data, and storage. In BAI, I was the infrastructure layer primarily focused around application performance, in standard north-south fashion. Packets in, packets out and my job was done. In AAI, I am the layer that decides whether your GPU cluster runs at 90 percent utilization or at 30 which indirectly makes me critical factor in ROI of AI investment. As per analysts, 70-90% of traffic inside a modern GPU cluster is now east-west, meaning GPU to GPU during training rather than user to server. A single AI cluster of 1000 servers carries up to 3.2 petabits per second of internal traffic during a training run. Every byte has to be lossless and latency sensitive. And your GPUs will still only hit 30 to 50 percent utilization, because the fabric, not the silicon, is where the ceiling actually sits.
While most people were watching GPUs, something drastic happened in networking space. Ethernet caught InfiniBand at the very top of the stack. The Ultra Ethernet Consortium shipped Spec 1.0 in June 2025 with more than a hundred member companies including NVIDIA itself. NVIDIA’s overall networking business, driven largely by Spectrum-X adoption in AI training fabrics, grew 162 percent year-over-year to $8.19 billion in Q3 fiscal 2026 alone. xAI’s Colossus supercomputer, running 555,000 GPUs, proved Ethernet at that scale by hitting 95 percent fabric throughput on Spectrum-X versus 60 percent on standard Ethernet.
Co-packaged optics crossed from research demo into production this year. Broadcom has shipped over 50,000 Tomahawk 5-Bailly CPO switches to Meta, with more than a million 400G port-hours of flap-free operation. NVIDIA Quantum-X Photonics is now shipping to customers in production. Marvell acquired Celestial AI for up to 5.5 billion dollars in December 2025 which highlights the shift happening from copper wires to optical connections to scale up bandwidth for massive AI clusters. [Every AI data center interconnect will be optical within five years, irrespective of distance.(https://nextwavesinsight.com/photonic-compute-production-lightmatter-ayar-labs/).
I am also joining Memory and Storage in supply-constrained territory. McKinsey projects 800G optical modules running 40 to 60 percent below demand through 2027, and 1.6T modules short by 30 to 40 percent through 2029. Indium phosphide laser supply is the current bottleneck, and there is no plausible path to scaling it quickly this decade.
To keep things interesting, AI training traffic now crosses continents to reach cheap power, which means the physical routes matter more than they used to. Red Sea subsea cables were cut again in 2026, disrupting connectivity between Europe and Asia for weeks. The Baltic saw five cable cuts between December 2024 and January 2025. Subsea cable investment is at its highest level in twenty years, because the industry finally noticed that where a packet physically travels is a strategic question.
Where I physically run matters, which country I traverse matters, and whether you own the rack the fabric terminates into now matters as well. All three of those questions used to be procurement details, and part of strategy meetings now.
The organizations that look smart in 2026 are the ones putting compute, memory, storage, and now network fabric physically close together, on infrastructure they own outright, running like cloud, with the operational simplicity that comes from treating infrastructure as a product instead of a project.
Here is what to walk away with from this confession. Your AI project’s ceiling is probably not the model or the GPU. It is probably me, running at 30 to 50 percent utilization because your fabric was designed for a different era and nobody in the building has the budget to rebuild it this quarter.
Next time someone tells you networking is a solved problem, ask them what the east-west traffic ratio inside their AI cluster looks like this month. Then notice how quickly the conversation ends after that.