Architecture to Autonomy Back to posts
Infrastructure & AI Economics

Hardware Is Eating AI,
and AI Is Eating Everything Else

Marc Andreessen said software would eat the world. It did. But in 2026, the economics have inverted. Physical infrastructure is the new bottleneck, and whoever controls compute controls the future.

By Samir Roshan Published April 12, 2026 20 min read
70%
Share of global DRAM production consumed by AI data centers in 2026
36 to 52wk
Lead times for data center GPUs in 2026
+171%
DRAM price increase driven by AI infrastructure demand
2028
Earliest projected year for meaningful supply relief
The A2A Pod: Listen Along
Hardware Is Eating AI, and AI Is Eating Everything Else

When Software Was Free and Scale Was Infinite

In 2011, Marc Andreessen published his famous essay arguing that software was eating the world. The thesis was simple: software has near-zero marginal cost. Once you write a piece of code, distributing it to a million users costs almost nothing. That economic reality gave software companies a structural advantage over every traditional industry. Retail, media, finance, transportation, healthcare. One by one, they fell.

And for fifteen years, the thesis held. Netflix displaced Blockbuster. Spotify reshaped music. Uber redrew transportation. Amazon swallowed retail. The common thread was that software's weightlessness, its freedom from atoms, was its superpower.

Software disrupted industries because it was cheap to copy and distribute. AI completely inverts that assumption. Every inference costs silicon, power, and water.

But something fundamental has changed. The next wave of technology, artificial intelligence, is not weightless. It is brutally, inescapably physical. Every model trained, every inference served, every reasoning chain computed demands GPUs, high-bandwidth memory, advanced packaging, electrical power, cooling infrastructure, and increasingly, water. The era of zero-marginal-cost disruption is over. Welcome to the era where atoms are eating bits.

Hardware Is Eating AI

If I had to write the 2026 version of Andreessen's essay, the headline would be: "Hardware is eating AI, and AI is eating everything else."

The logic is simple. AI is doing exactly what software did, disrupting every industry it touches. But unlike software, AI cannot scale for free. Its supply chain has chokepoints that look more like the oil industry than the software industry. The companies that control scarce hardware hold enormous power, almost like OPEC did with crude.

Look at the bottleneck stack. It is not just GPUs. The binding constraint is actually memory and packaging:

01
HBM Memory
SK Hynix, Samsung, and Micron produce 90% of the world's memory chips. Production has been structurally redirected toward AI, away from consumer devices.
02
CoWoS Packaging
TSMC's Chip-on-Wafer-on-Substrate process bonds HBM to GPU dies. It is the single biggest bottleneck, fully allocated through mid-2027.
03
GPU Allocation
Microsoft, Google, Meta, and Amazon placed multi-billion-dollar forward orders consuming most NVIDIA allocation through 2027, crowding out everyone else.
04
Energy & Cooling
A single NVIDIA B200 draws up to 1,000 watts. Data centers are maxing out power grids, with some integrating dedicated natural gas plants.

All of this creates a cascading supply crisis. NVIDIA has reportedly cut consumer RTX 5000-series GPU production by 30 to 40% to prioritize data center chips. Micron discontinued its entire consumer Crucial memory lineup to focus on AI. PC prices are climbing. Even game developers are being forced into unplanned optimization work because RAM costs have restructured what is affordable to ship.

This is not a GPU shortage. It is a memory shortage, a packaging shortage, a power shortage, and a supply chain architecture problem, all compounding at once.

What Do the Projections Actually Say?

The question everyone is asking: is this a structural shift or a cyclical bottleneck? The data suggests it is cyclical, but with a longer tail than most people expect.

Supply-side timeline

2025 to Mid 2026
Peak Constraint
HBM and CoWoS capacity fully allocated. GPU lead times stretch to 36 to 52 weeks. Hyperscalers consume most available supply. Consumer hardware markets absorb collateral damage through price hikes and production cuts.
Late 2026 to 2027
Gradual Stabilization
New fabs from Intel, TSMC, Samsung, and Micron begin ramping. HBM and DDR5 production improves. But because AI demand is also growing, supply expansion may only meet, rather than exceed, consumption. Prices remain above historical norms.
2028+
Meaningful Relief
New chip factories (which take 3 to 5 years to build) reach production maturity. Alternative accelerators capture meaningful market share. Model efficiency gains reduce per-inference hardware requirements. The cost curve begins bending back toward software-like economics.

DRAM consumption by AI: the elephant in the room

One number tells the whole story: AI data centers are projected to consume approximately 70% of all global DRAM production in 2026. That single statistic explains why everything from laptops to gaming consoles to Raspberry Pis is getting more expensive. The memory market has been structurally redirected.

2023
~25% AI share of DRAM
2024
~40% AI share of DRAM
2025
~55% AI share of DRAM
2026
~70% AI share of DRAM (projected)

TrendForce researchers have described this reallocation of memory capacity toward AI as "permanent." Even when new capacity comes online, a significant share will remain tied to AI customers. This is not a temporary diversion. It is a structural rebalancing of who gets to buy memory and at what price.

Five Forces That Will Break the Bottleneck

If this were purely about physics, finite silicon, finite fabs, the shortage would be permanent. But technology markets correct themselves: when something is expensive enough, the entire industry innovates around it. The relief comes from five directions.

1. Alternative accelerators and inference-specific silicon

The GPU monoculture is ending. Google's TPU v6e already delivers up to four times better performance per dollar for inference compared to NVIDIA's H100. Amazon's Trainium, Intel's Gaudi, and various FPGA-based solutions are diversifying the compute supply chain. XPU spending is projected to grow 22% in 2026, and by 2027, these alternatives may account for a larger share of AI hardware budgets than GPUs.

But the most telling signal came on Christmas Eve 2025, when NVIDIA itself paid $20 billion to license Groq's Language Processing Unit (LPU) technology. Think about what that means: the company that dominates GPUs spent its largest deal ever to acquire an architecture that does not use GPUs at all.

Groq's LPU completely bypasses the HBM memory bottleneck that sits at the center of the 2026 shortage. Instead of using off-chip HBM (the scarce, expensive memory that SK Hynix, Samsung, and Micron cannot produce fast enough), Groq built its chips with on-chip SRAM as primary weight storage. SRAM is roughly 100 times faster than HBM and 20 times more energy-efficient per bit of data access. The trade-off is capacity: each LPU chip holds only about 230 to 500 MB, so you need hundreds of chips linked together to run a large model. But those chips do not consume any HBM, do not require TSMC's constrained CoWoS packaging, can run on older 14nm fabrication nodes, and are air-cooled with no liquid cooling infrastructure needed.

The inference efficiency numbers make the strategic logic clear. NVIDIA's own published figures for the Groq 3 LPU show approximately 150 tokens per watt, compared to roughly 4.3 tokens per watt for the H100 in comparable model serving. That is a 35x improvement in inference throughput per megawatt. For data center operators running millions of inference requests daily, that changes the economics entirely.

NVIDIA's $20 billion Groq deal is an admission: the GPU architecture that dominates AI training is not the optimal answer for inference. And inference now accounts for over 55% of all AI compute.

This matters for the shortage because AI is shifting from a training-dominated phase to an inference-dominated phase. Deloitte estimates inference workloads now account for roughly two-thirds of all AI compute, up from one-third in 2023. The market for inference-optimized chips alone is projected to exceed $50 billion in 2026. If a meaningful share of those inference workloads can move to architectures that do not consume HBM, CoWoS slots, or cutting-edge fab capacity, it directly relieves pressure on the entire supply chain.

The point extends beyond Groq. Arm launched its AGI CPU in March 2026, designed specifically for the decode phase of AI inference and agentic AI orchestration. Arm's CEO predicted the chip alone would generate $15 billion in revenue by 2031, and forecast that emerging inference workloads would quadruple demand for CPUs. Intel and SambaNova announced a multi-year collaboration built on a modular architecture where GPUs handle prefill, SambaNova's RDUs handle decode, and Xeon CPUs orchestrate the system. Microsoft shipped Maia 200, its own 3nm inference accelerator, already deployed in Azure powering Copilot and GPT workloads. Cerebras continues to push wafer-scale SRAM-based inference at speeds six times faster than Groq on comparable models.

Not every AI workload needs a top-tier GPU with HBM. Many production inference tasks, particularly in enterprise settings, can run efficiently on CPUs, specialized ASICs, or lighter accelerators. The industry is waking up to the fact that the right chip for the job is no longer always a GPU, and that realization is itself a release valve on the shortage.

2. Model efficiency is improving fast

The era of "just throw more GPUs at it" is giving way to smarter engineering. Smaller, distilled models that run on a fraction of the hardware are becoming viable for most production workloads. Techniques like quantization, mixture-of-experts, and speculative decoding are reducing the compute cost per inference by orders of magnitude.

3. Processing-in-memory will change the game

The real bottleneck in 2026 AI is not computation. It is memory bandwidth. The cost of shuttling data between memory and processor is where most energy and time is consumed. Samsung and SK Hynix are investing heavily in Processing-in-Memory (PIM) architectures, where computation happens directly inside the memory chips. This could significantly reduce the data movement problem that drives current hardware costs.

4. New fabs are coming online

The capital is flowing. Intel, TSMC, Samsung, and Micron all have major fab construction projects underway. The challenge is physics: chip factories take 3 to 5 years to build, equip, and certify. But by 2028, the supply side will look very different from today.

5. Neuromorphic and novel architectures

The human brain runs on about 20 watts of power. Current AI clusters use megawatts. That gap represents one of the largest efficiency opportunities in the history of computing. Research into neuromorphic chips, processors that mimic the brain's use of spikes and pulses rather than continuous electricity, is accelerating because the current power economics are unsustainable.

Custom Silicon and Software-Defined Efficiency: Broadcom's Two-Pronged Answer

One company is uniquely positioned to attack the hardware crisis from both sides of the equation: Broadcom. On one hand, it is building the custom silicon that breaks the GPU monoculture. On the other, its VMware Cloud Foundation platform helps enterprises squeeze maximum value from the hardware they already own. That combination makes Broadcom one of the most important players in the AI infrastructure story right now.

The TPU partnership: custom silicon at scale

On April 6, 2026, Broadcom filed an 8-K disclosing a long-term agreement with Google to develop and supply future generations of custom Tensor Processing Units through 2031. This is not a chip order. It is a five-year infrastructure commitment covering processors, networking fabric, and the entire AI rack architecture.

The deal also pulled in Anthropic. Under the agreement, Anthropic will access approximately 3.5 gigawatts of next-generation TPU-based compute capacity starting in 2027. To put that in perspective, 3.5 gigawatts rivals the energy output of multiple nuclear power plants. The build-out is estimated to cost between $120 billion and $175 billion, making it one of the largest infrastructure investments in technological history.

The complexity of AI workloads has outpaced the capabilities of general-purpose silicon. Custom-tailored XPUs are the only way to achieve the power efficiency required for the next million-node AI clusters.

Why does this matter for the hardware crisis? Because custom TPUs deliver better performance per watt than general-purpose GPUs for specific workloads. Google's latest Ironwood TPU v7, built on a 3-nanometer process, offers a four-fold performance improvement over previous generations while consuming less power per teraflop. When you can do more work per chip, you need fewer chips. When you need fewer chips, you ease the pressure on memory, packaging, and power.

Broadcom's AI revenue trajectory tells the story. From $4.4 billion in Q2 FY2025 to $8.4 billion in Q1 FY2026, that is 106% year-over-year growth. The company is targeting $100 billion in AI chip revenue by 2027, with a $73 billion backlog already in hand. Broadcom has captured 60 to 70% of the custom AI accelerator market, and its ability to integrate custom silicon with high-speed networking (Tomahawk and Jericho Ethernet switches, industry-leading SerDes IP) creates a full-stack offering that competitors cannot easily replicate.

Anthropic's approach here is worth noting. The company is explicitly taking a multi-hardware strategy, using AWS Trainium, Google TPUs, and NVIDIA GPUs in parallel. The market is not consolidating around one chip. It is fragmenting into strategic, workload-specific choices. That fragmentation is itself a release valve on the shortage.

VCF 9.0: do more with the hardware you already have

The other half of Broadcom's answer comes from VMware Cloud Foundation. While the TPU partnership addresses the supply side (build better chips), VCF attacks the demand side (use existing hardware more efficiently).

Broadcom's VCF division has been direct about this. Krish Prasad, SVP and GM of the VMware Cloud Foundation Division, said it plainly: the answer to a hardware crisis is not more hardware. It is smarter software. VCF 9.0 was engineered with the 2026 supply crunch in mind, built around four pillars of efficiency.

01
NVMe Memory Tiering
VCF 9.0 uses NVMe drives as a secondary memory tier, doubling host memory capacity at a fraction of DRAM cost. With DDR5 at $40/GB and NVMe at $1/GB, that is a 40:1 cost ratio. Up to 42% lower memory and server TCO.
02
Hardware Elimination
By virtualizing load balancing (VMware Avi) and security (VMware vDefend), VCF removes the need for expensive, power-hungry proprietary hardware appliances entirely.
03
AI-Native Platform
Private AI Services are now standard in VCF 9.0: GPU monitoring, model storage, model runtime, agent builders, vector databases, and data indexing. All built in, no bolt-ons needed.
04
Multi-Accelerator Support
VCF supports both NVIDIA Blackwell GPUs and AMD Instinct MI350 Series without requiring application rewrites. This gives enterprises the freedom to use whichever accelerator they can actually source.

NVMe memory tiering deserves special attention because it directly addresses the memory crisis at the heart of the 2026 shortage. The way it works: ESXi intelligently tracks which memory pages are actively being used ("hot") and which are sitting idle ("cold"). Hot pages stay in DRAM for full-speed access. Cold pages get moved to NVMe storage, which is slower but orders of magnitude cheaper. The VM sees a single, larger memory space. It does not know the difference.

The economics are striking. A host with 1TB of DRAM can now present 2TB of total memory at the default 1:1 ratio, and up to 4TB at a 1:4 ratio for workloads with low active memory like VDI. Broadcom's own testing shows less than 5% performance loss when comparing 1TB DRAM-only to 1TB memory tiering configurations. For database workloads, they were able to double VM density per host with minimal performance impact. On a Dell PowerEdge R760, the per-server cost dropped from $55,878 to $33,792, a savings of nearly 40%. In a world where DRAM prices have surged 95% and memory now accounts for over half the cost of a new server, that is not a marginal improvement. It changes the procurement math entirely.

The feature also integrates natively with vMotion, DRS, HA, and vSAN. No special workflows. Most servers already certified for vSphere 8.x carry forward into VCF 9.0, so organizations can start using NVMe tiering on their existing fleet without waiting for new hardware that may take quarters to arrive.

The broader TCO impact compounds from there. Broadcom's VMware Telco Cloud Platform 9, built on VCF 9.0, claims 40% cumulative TCO savings over five years compared to siloed architectures. When server DRAM prices have surged nearly 95% and lead times stretch past a year, that kind of efficiency gain is not an optimization exercise. It is a strategic lifeline.

Nine of the top ten Fortune 500 companies have committed to VCF, with over 100 million cores licensed worldwide. The platform's appeal in this moment is clear: when you cannot buy your way out of a hardware shortage, you need software that makes every rack, every GPU, every gigabyte of memory count for more.

Broadcom is playing both sides of the board. On the silicon side, it is building the custom chips that offer a credible alternative to the NVIDIA GPU monoculture. On the platform side, it is giving enterprises the tools to stretch their existing infrastructure further. That dual position, chip architect and software platform, is unique in the industry. And in a hardware-constrained world, it may be the most valuable position of all.

The Haves and the Have-Nots

The most consequential near-term effect of the hardware shortage is a split in the AI world:

The Haves

  • Hyperscalers with multi-billion-dollar GPU commitments
  • Access intelligence via cloud APIs at scale
  • Can afford 36 to 52 week lead times through forward contracts
  • Building custom silicon (Google TPU, Amazon Trainium)
Compute Divide

The Have-Nots

  • Startups and mid-market companies priced out of GPU allocation
  • Run smaller, distilled models on local hardware
  • On-demand cloud pricing at 2 to 3x premium, often throttled
  • Forced into creative multi-chip and edge strategies

This is the dynamic that should concern strategists most. The AI revolution promises democratization of intelligence, but its infrastructure economics are pushing toward concentration. The companies that locked in forward GPU contracts in 2024 and 2025 have a multi-year structural advantage over everyone who did not. That advantage compounds. More compute means better models, which means more customers, which means more revenue to reinvest in compute.

Silicon is the new oil. Whoever controls the compute supply chain controls the AI economy. The question is how long that oligopoly lasts.

What This Means for Enterprise Leaders

If you are an enterprise architect, CTO, or infrastructure leader, the hardware shortage demands a different kind of strategic thinking than the software era required. Software strategy was about talent and speed. AI strategy is about supply chain resilience.

01
Diversify Compute
Do not bet everything on NVIDIA. Evaluate TPUs, Trainium, and FPGAs. Companies implementing multi-chip strategies are seeing 30 to 50% lower total cost of ownership.
02
Invest in Efficiency
Model optimization is not a nice-to-have. It is a survival strategy. Quantization, distillation, and right-sizing models to workloads can reduce hardware needs by an order of magnitude.
03
Plan Procurement in Years
The era of just-in-time GPU purchasing is over. Lead times are measured in quarters, not weeks. Financial hedging and long-term contracts are now standard practice.
04
Consider Geography
Electricity costs vary by 3x across regions. Eastern European data center locations offer 25% lower construction costs. Energy strategy is now infrastructure strategy.

Cyclical, Not Structural. But Respect the Cycle

History suggests the hardware shortage is not permanent. Every major technology bottleneck, from DRAM in the 1990s to hard drives after the Thailand floods to chip shortages during COVID, eventually resolved through a combination of investment, innovation, and demand moderation.

The AI hardware crunch will resolve too. Capital is flooding into new fabs. Alternative architectures are gaining traction. Model efficiency is improving faster than model size is growing. By 2028, the supply picture will look very different.

But "not permanent" does not mean "not consequential." The companies and strategies that are forged in this 2025 to 2028 window of constraint will define the competitive picture for a decade. Those who secure compute now, diversify their chip strategies, and invest in efficiency will have compounding advantages. Those who wait for the market to normalize will find that their competitors used the bottleneck as a forcing function to build better, leaner, more resilient infrastructure.

Software ate the world because it was free to distribute. AI is eating the world despite being expensive to run. That says something about how powerful the technology is. But it also means the winners of the AI era will not just be the ones who build the best models. They will be the ones who solve the hardware equation.

The Bottom Line

The narrative has shifted from "software eats the world" to "hardware eats AI, and AI eats everything else." The shortage is cyclical, not structural. Relief arrives by 2028. But the strategic choices made during this window of scarcity will define competitive advantage for years beyond it. Treat compute supply chain as a first-class strategic capability, not a procurement function.

Views expressed here are my own and do not represent the views of any current or former employer, client, or affiliated organization.