Every enterprise AI conversation I sit in narrows to the same question within ten minutes: which model should we standardize on? Fair question, wrong place to start. It sits so far downstream that it’s like choosing the paint before checking the plot for water and power. The models everyone names, ChatGPT, Gemini, Claude, DeepSeek, are the part you see and click on. That’s why they get the attention and the board deck.
Here is what the hype misses. Judging AI by its models is like judging the car industry by the showroom. Behind that front sit foundries and supply chains that took decades and hundreds of billions to build, none turning over on a quarterly roadmap. The fab that made its chips took five years and tens of billions; the machine inside it took twenty years to invent. Software time against infrastructure time: that is the whole story, and it breaks into five layers most people running a P&L never think about.²,⁷ And there’s a clock. The EU AI Act wants providers of high-risk systems to name the suppliers underneath them—duties Brussels shoved to December 2027 in this year’s Digital Omnibus. That bought room, not a pardon. Dependency mapping is paperwork now.⁹
Layer 1: Compute
Everything starts here: every model resolves into arithmetic run on chips. For now that means GPUs, gaming hardware that turned out to be the ideal engine for training neural networks. NVIDIA rode that accident to one of the largest market caps in history; AMD is grinding to catch up. Worn down by NVIDIA’s margins and waiting list, the hyperscalers now roll their own: Google’s TPUs, Amazon’s Trainium and Inferentia, Microsoft’s Maia, Huawei’s Ascend, the last a centerpiece of China’s push not to need anyone else.¹
Here is what that roster hides. The chip is the easy half; NVIDIA’s real moat isn’t the silicon, it’s CUDA, nearly two decades of libraries and kernels almost every model and framework is written against. So when a cheaper accelerator appears, switching isn’t a purchase order, it’s a porting project: rewrite the kernels, re-tune performance, chase the regressions. That is why they pour fortunes into custom silicon and still buy NVIDIA by the crate. Compute is now a strategic input on the order of capital. You lock it in years ahead, or watch your roadmap slip while a better-provisioned rival ships.
Layer 2: Manufacturing
Here is the catch: securing compute assumes somebody can build the chips, and almost nobody can. A handful of firms manufacture at the leading edge, forming one interdependent chain. TSMC fabricates a huge share of the most advanced parts. SK hynix and Samsung supply the high-bandwidth memory those accelerators need. Upstream of it all sits ASML, in the Dutch town of Veldhoven, the only company on Earth that can build an EUV lithography machine.
That machine borders on science fiction, and took ASML two decades to make work: it fires a laser at droplets of molten tin tens of thousands of times a second, bursting each into plasma to make the light that prints the smallest features. Older kit still prints fine features. Multi-patterning does it commercially—just slower, pricier, more passes. So EUV wins on the spreadsheet long before physics.⁴,⁵,⁶ And these aren’t swappable: roughly $220 million a unit, past $380 million for the newest High-NA generation, each 150 tonnes, shipped in 250 crates, a crew the better part of a year to stand up.
Now stack a second fragility: nearly all leading-edge output happens on one island. TSMC is expanding into Arizona and Japan, but the frontier still runs through Taiwan, on a fault line and a geopolitical flashpoint. If that chain snaps, new accelerators stop arriving, cloud GPU capacity dries up, prices spike, and every roadmap assuming “more compute next year” freezes. Washington bars ASML from selling EUV to China at all, which is why China’s advanced fabs improvise around older tools. No single country owns the chain end to end. A few can pinch it shut.
Layer 3: Cloud
Look, the manufacturing chokepoint is at least visible. The dependency that creeps up on you is one you walk into willingly. Hardly anyone “buys AI software” now; they plug into a platform, Azure, AWS, Google Cloud, Oracle, that arrives with compute, models, storage, security, and orchestration already integrated, and entangled. The model is one part of the bundle, and the bundle is the product.
The trap has a name: data gravity. Data attracts services and more data, and the pile compounds. Getting it in is cheap by design; getting it out is not. Providers charge egress fees to move data off their cloud while charging nothing to bring it in. Your queues, identity model, observability, and legacy middleware are wired to provider-specific services, so leaving means refactoring the plumbing, not copying files, and re-earning every compliance certification you hold in the new vendor’s regions. Standardize on a stack and your AI roadmap becomes a function of whatever that vendor ships next. Switching costs were never a contract clause. They are the accumulated weight of every integration you built. Embed AI deeply enough and the vendor you picked for speed becomes the platform you architect around for a decade.²,⁸ That is an architecture decision wearing procurement’s clothing.
Layer 4: Energy
Power is the layer almost nobody wrote into the business case. Training frontier models consumes enormous amounts of electricity, and inference only climbs as adoption spreads. The IEA’s 2025 numbers are hard to wave off: data-centre demand is tracking toward roughly 945 terawatt-hours a year by 2030, about what all of Japan consumes today, and in the US, data centres are on course to draw more power than aluminium, steel, cement, and chemicals production combined.³
Here is where it stops being an engineering line item and becomes local politics. That power comes off a grid that already serves factories and homes. Ask for a gigawatt in some county and you are bidding against the manufacturer that wanted to expand and the neighborhood worried about its bills. Interconnection queues stretch for years; the IEA reckons roughly a fifth of planned data-centre projects risk delay for grid reasons alone, and in some towns data centres already pull the lion’s share of local electricity. That is why more communities are placing moratoriums on new builds, and why Microsoft, Google, Amazon and Meta now behave less like software companies than utilities: locking up renewables, signing with nuclear operators, firing up on-site gas because the grid cannot connect them fast enough.³ Cheap, reliable power is now a siting constraint: it decides where the next capacity can go, and where it can’t.³,⁷
Layer 5: Geopolitics
Forget the scoreboard. This is a chokepoint map—not the US-versus-China cage match the headlines keep selling. The US holds models, money and cloud. China holds the inputs, and is spending its way toward chips it still can’t build. Taiwan holds fabrication—for exactly as long as TSMC keeps doing what nobody else can, a thought that should keep your risk officer awake. Europe holds ASML, plus the rulebook everyone obeys anyway. India and the Gulf are buying in, talent and sovereign-funded campuses.⁴,⁵,⁶,⁷ The real chokehold hides in boring inputs: China refines something like 99% of the world’s gallium as of 2025, a direct input to advanced chips, so one country’s export desk can send a tremor through everyone else’s supply chain.³ Nobody wins outright. Just rungs—and every rung is somebody’s grip.
Where the Advantage Lives
So the useful reframe isn’t “which model.” It’s a cluster of harder questions. Whose ecosystem are we joining. How fragile is the chain beneath us. What breaks us if the politics shift. And the one that counts: where do we create value that isn’t trivially replaceable? Those aren’t IT questions so much as strategy questions with IT inside them, which is why they belong in front of the board, not three levels down in an architecture review.²,⁷,⁸
In practice, that means an honest audit. Map every layer you depend on and name the single points of failure: one chip vendor, one cloud, one fab, one region. Stress-test the exit for each. Get concrete. If your models sit on one hyperscaler’s managed endpoints, price the escape—engineer-months, re-certification, egress—and put that number in the risk register before the renewal, not after. If leaving takes eighteen months and a fortune, that isn’t a supplier, it’s a dependency. Then find the durable edge that a better model next quarter can’t copy. It is almost never the model. It is proprietary data, a regulated workflow you earned the right to run, distribution, hard-won trust, institutional knowledge that doesn’t port. The models still matter, but the winners over the next decade won’t hold a marginally better one; that edge evaporates inside a release cycle. They will be the ones who read the layers underneath early enough to position on purpose, and built something hard to take away.
References
¹ NVIDIA Corporation, “NVIDIA Blackwell Platform and AI Infrastructure”, NVIDIA https://www.nvidia.com accessed 24 July 2026
² McKinsey & Company, “The Economic Potential of Generative AI: The Next Productivity Frontier”, McKinsey Global Institute, 2023 https://www.mckinsey.com accessed 24 July 2026
³ International Energy Agency, Energy and AI (IEA 2025) https://www.iea.org accessed 24 July 2026
⁴ Semiconductor Industry Association, “Powering AI: The Semiconductor Ecosystem at the Foundation of Data Centers”, SIA, 2026 https://www.semiconductors.org accessed 24 July 2026
⁵ TSMC, Annual Report 2025 (Taiwan Semiconductor Manufacturing Company 2025) https://www.tsmc.com accessed 24 July 2026
⁶ ASML, Annual Report 2025 (ASML Holding NV 2025) https://www.asml.com accessed 24 July 2026
⁷ World Economic Forum, “Artificial Intelligence and the Future of Global Value Chains”, WEF https://www.weforum.org accessed 24 July 2026
⁸ OECD, “OECD Artificial Intelligence Policy Observatory”, OECD https://oecd.ai accessed 24 July 2026
⁹ Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 (Artificial Intelligence Act), arts 25–26 and annex IV, as amended by the Digital Omnibus on AI (2026) https://eur-lex.europa.eu accessed 24 July 2026
