NVIDIA
Designs the GPUs and accelerated computing platforms used for much of modern AI training and inference, and sells the networking that ties them together.
Who builds what, and why it matters
Every city needs foundations, power, roads, buildings, and people using them. AI is the same, just faster and more expensive.
Written by Matthew Bernath
Seven core layers. Two frontier theses. One map of where AI's money actually goes.
Read the ecosystem broadly from infrastructure to application. The layers depend on and reinforce one another, but the real architecture is not perfectly linear. Security, governance and demand are not steps in the sequence: they cut across everything, and they get their own treatment below.
Read it bottom to top. Each layer only exists because the ones beneath it do. Security and governance are not steps in the sequence: they wrap the whole thing.
A request travels up the stack. Value flows back down as revenue.
Pick a path and see which layers a real question, a real dollar, or a real product actually touches.
A prompt typed into an app
Adds context, permissions and tools
Turns the prompt into tokens and a response
Retrieves the records the answer depends on
Schedules the work onto real machines
Moves data between processors and sites
Does the arithmetic, fed by memory bandwidth
Pays for it in joules and removes the heat
Real architecture is not perfectly linear. Steps run in parallel, repeat, and sometimes get skipped entirely by caching.
AI can run on conventional processors, but specialised accelerators make large modern AI workloads dramatically more efficient. This layer designs, manufactures and packages them.
Tap or hover a box to see what it does
CPUs are optimised for flexible, general purpose and low latency computing. GPUs contain many more, simpler processing units and are especially effective at performing large numbers of similar calculations in parallel, which is what training and inference mostly consist of.
Designs the GPUs and accelerated computing platforms used for much of modern AI training and inference, and sells the networking that ties them together.
Designs CPUs and the Instinct family of AI accelerators, sold as an alternative to NVIDIA at the data centre scale.
Designs and manufactures CPUs, and is attempting to build a contract manufacturing business for other companies' chip designs.
Designers such as NVIDIA and AMD are fabless: they own no factories. The physical work happens at foundries, using lithography systems that print features measured in nanometres, and packaging that binds logic and memory into a single module.
Manufactures chips designed by other companies, including most leading edge AI accelerators, and provides the advanced packaging that binds logic dies to memory.
Builds the lithography systems used to print circuit patterns onto silicon wafers, including extreme ultraviolet machines.
HBM stacks memory dies vertically and places them beside the accelerator, giving far more bandwidth than conventional memory laid out on a board. Since large models must stream billions of parameters for every token they produce, HBM supply has repeatedly gated how many AI systems can actually be built.
Manufactures memory, including high bandwidth memory stacks that sit alongside AI accelerators.
Manufactures DRAM, NAND and high bandwidth memory used in AI servers.
Manufactures memory including HBM, runs a contract chip manufacturing business, and builds consumer devices.
Google's TPU, Amazon's Trainium and Inferentia, Microsoft's Maia and Meta's MTIA exist because the largest buyers would rather own their cost structure than rent it. Designing in-house improves margins on their own workloads and reduces dependence on a single supplier. For NVIDIA this is both an opportunity, since it still sells into those same data centres, and a risk, since the customers designing these parts are also its biggest source of revenue.
Co-designs custom AI accelerators for hyperscalers and supplies much of the switching and connectivity silicon inside data centres.
Builds Gemini models through Google DeepMind, designs TPU accelerators, operates Google Cloud, and distributes AI through Search, Workspace and Android.
Runs AWS, designs Trainium and Inferentia accelerators, hosts third party models through Bedrock, and invests in Anthropic.
Runs Azure, distributes AI through Microsoft 365 and Copilot, partners commercially with OpenAI, and designs its own Maia accelerators.
Develops the Llama model family and deploys AI across its own products, while designing MTIA accelerators for internal workloads.
Designs data infrastructure silicon including optical interconnect, custom compute and storage controllers.
AI's infrastructure problem is increasingly not just whether GPUs can be bought, but whether they can be powered and cooled. That runs from generation through the grid to the rack.
Tap or hover a box to see what it does
Nuclear, gas, renewables, fuel cells and on site distributed generation. Buyers increasingly want power that is both firm and low carbon, which is a harder combination to source than either on its own.
Operates a large fleet of nuclear generation in the United States and sells power under long term contracts.
Runs a mixed generation fleet including nuclear, gas and storage, and sells into wholesale power markets.
Develops and operates renewable generation and storage at scale alongside a regulated Florida utility.
Makes solid oxide fuel cells that generate electricity on site from natural gas, biogas or hydrogen.
Designing small modular fast reactors intended to sell power directly to large users such as data centres.
Transmission lines, substations, transformers and the connection agreements that determine whether a site can draw the power it has contracted for. This is where most of the waiting happens.
Supplies gas turbines, grid equipment and wind generation, and services the installed base.
Supplies electrification and automation equipment including transformers, switchgear and drives.
Uninterruptible power supplies, switchgear, busway and backup systems sit between the substation and the rack. Availability is engineered here rather than assumed, and no design eliminates outage risk entirely.
Supplies electrical distribution equipment, switchgear, busway and uninterruptible power systems.
Supplies power distribution, UPS, cooling and data centre management software, often as a packaged design.
Air cooling struggles as rack densities climb, so liquid cooling has moved from specialist to routine at the high end. Whatever the method, the heat still has to be rejected somewhere, which is why water and site selection keep appearing in the conversation.
Makes data centre power and thermal management equipment, including liquid cooling for high density racks.
Makes thermal management products including data centre cooling and heat rejection systems.
The buildings, campuses and cloud platforms where AI actually runs. Hyperscale cloud, colocation, specialised AI and HPC facilities, and sovereign infrastructure are four different businesses.
Tap or hover a box to see what it does
The largest platforms sell compute, storage and managed AI services globally, organised into regions and availability zones so that a failure in one place does not take the service with it. Their own AI campuses are increasingly described as AI factories: purpose built sites optimised for training and serving models rather than for general computing.
Runs Azure, distributes AI through Microsoft 365 and Copilot, partners commercially with OpenAI, and designs its own Maia accelerators.
Runs AWS, designs Trainium and Inferentia accelerators, hosts third party models through Bedrock, and invests in Anthropic.
Builds Gemini models through Google DeepMind, designs TPU accelerators, operates Google Cloud, and distributes AI through Search, Workspace and Android.
Runs Oracle Cloud Infrastructure with a strong GPU cluster business, and sells the databases and applications that hold a great deal of enterprise data.
Customers place their own equipment in someone else's building and connect to networks, clouds and each other. Power density per rack is the number that decides whether a given facility can host modern AI hardware at all.
Operates colocation data centres where customers place their own equipment and interconnect with each other.
Develops and leases large scale data centre capacity to hyperscalers and enterprises.
Operators built specifically around dense GPU clusters, high speed interconnect and the workloads that need them. Sovereign AI infrastructure is a related category, driven by rules about where data and models may physically reside.
Operates GPU cloud infrastructure purpose built for AI training and inference rather than general computing.
Operates AI focused cloud infrastructure in Europe. It emerged from the old Yandex international business after the Russian operations were sold in 2024, with the Dutch parent keeping the global assets.
Bitcoin mining, and historically Ethereum mining, required many of the same ingredients now sought by AI infrastructure developers: large power allocations, suitable sites and industrial cooling. Ethereum itself moved from proof of work to proof of stake in September 2022, so it is no longer mined.
Training a large model means tens of thousands of processors behaving as one machine. Networking can become the bottleneck even when there is plenty of compute available.
Tap or hover a box to see what it does
The shortest and fastest hop. NVIDIA's NVLink and NVSwitch connect accelerators inside a server so that several GPUs behave as one larger device, with far more bandwidth than a standard network link could carry.
Designs the GPUs and accelerated computing platforms used for much of modern AI training and inference, and sells the networking that ties them together.
Across a rack and a hall, clusters run on InfiniBand or high speed Ethernet. NVIDIA supplies InfiniBand and its Spectrum-X Ethernet platform following the Mellanox acquisition, while Arista and Broadcom drive the open Ethernet alternative and Credo supplies the high speed connectivity between them.
Designs the GPUs and accelerated computing platforms used for much of modern AI training and inference, and sells the networking that ties them together.
Builds high speed Ethernet switching and the network operating system used in large data centres and AI clusters.
Co-designs custom AI accelerators for hyperscalers and supplies much of the switching and connectivity silicon inside data centres.
Makes high speed connectivity products including active electrical cables and retimers used inside racks.
Beyond the hall, traffic moves over optics. Transceivers, lasers and photonic components carry campus and long haul links, and co-packaged optics is the direction of travel as data rates climb.
Makes optical transceivers, lasers and photonic components used to move data between racks and buildings.
Supplies photonics including lasers and transceivers for data centre interconnect.
Designs data infrastructure silicon including optical interconnect, custom compute and storage controllers.
A pre-revenue research company developing polymer based electro-optic materials for high speed optical modulation.
Between regions, and then to the people and devices actually making requests. Latency at this distance shapes where inference is served from rather than where models are trained.
Operates a global network providing content delivery, DDoS protection, zero trust services and edge compute.
The city's libraries, records and archives. Models are only one part of the system: enterprises also need data that is accessible, trusted, governed, structured, searchable, permissioned and connected to business context.
Tap or hover a box to see what it does
Where analytical data is consolidated, modelled and governed. The lakehouse pattern merged the cheap storage of a data lake with the governance and query behaviour of a warehouse, mostly on open table formats.
Runs a cloud data platform for storing, governing and querying enterprise data, with AI features layered on top.
Provides a lakehouse platform combining data engineering, analytics and machine learning on open table formats.
Privately held. Investors include Microsoft and several large asset managers, so exposure is indirect and partial.
The systems that run the business in real time, plus the pipelines that move events between them. Agents acting on yesterday's data cause expensive mistakes, so freshness is an architectural requirement rather than a nicety.
Provides a document database, delivered mainly as the Atlas managed service, with integrated vector search.
Runs Oracle Cloud Infrastructure with a strong GPU cluster business, and sells the databases and applications that hold a great deal of enterprise data.
Commercialises Apache Kafka as a managed platform for streaming data between systems.
Retrieval turns a general model into one that knows about your business. Vector search finds records by meaning rather than exact wording, and metadata, lineage and permissions decide what may be returned to whom.
Provides search, observability and security analytics built on Elasticsearch, including vector search.
Sells data integration and decision platforms to governments and large enterprises, increasingly packaged around AI driven workflows.
Provides enterprise resource planning software that runs core finance, supply chain and HR processes for many large firms.
The organisations training foundation models and serving them through APIs and products. Several of the most important are private companies, which is inconvenient but does not make them less important.
Tap or hover a box to see what it does
Organisations training the largest general purpose models. Several are private, so the practical question for investors is usually which listed business carries meaningful exposure to them.
Develops the GPT family of models and ships them through ChatGPT, an API and enterprise products.
Private. Microsoft is a major shareholder and its principal cloud and commercial partner, which is the most common route to indirect exposure.
Develops the Claude family of models, sold through an API, consumer apps and cloud marketplaces.
Private. Amazon and Alphabet are both significant investors and cloud partners, which is the usual indirect route.
Alphabet's AI research organisation, responsible for the Gemini model family and a long research record beyond language models.
Develops the Grok model family and operates its own large training cluster.
Private, with reported links to other Musk controlled businesses. There is no clean listed proxy.
Model families whose weights can be downloaded and run independently. That is not the same as open source: licences often restrict use, and the Open Source Initiative's definition requires more than published weights.
Develops the Llama model family and deploys AI across its own products, while designing MTIA accelerators for internal workloads.
Develops efficient models, several released with open weights, aimed at European enterprises and sovereign deployments.
Private, French, with strategic investors including industrial and technology partners. Exposure is largely unavailable to public investors.
How models actually reach customers. Cloud marketplaces, enterprise agreements and productivity software are the routes through which most organisations first pay for AI.
Runs Azure, distributes AI through Microsoft 365 and Copilot, partners commercially with OpenAI, and designs its own Maia accelerators.
Runs AWS, designs Trainium and Inferentia accelerators, hosts third party models through Bedrock, and invests in Anthropic.
Builds Gemini models through Google DeepMind, designs TPU accelerators, operates Google Cloud, and distributes AI through Search, Workspace and Android.
Designs the GPUs and accelerated computing platforms used for much of modern AI training and inference, and sells the networking that ties them together.
Companies that use AI to deliver products, workflows and services people actually buy. This is where model capability meets a budget holder with a problem.
Tap or hover a box to see what it does
Software that runs institutional processes, with AI increasingly embedded as agents that take actions rather than only answer questions. Buying criteria here are auditability and integration long before they are model benchmarks.
Sells data integration and decision platforms to governments and large enterprises, increasingly packaged around AI driven workflows.
Provides a workflow platform for IT, HR and customer operations, with AI agents embedded in those workflows.
Sells customer relationship management software with an agent platform layered over its data and workflow estate.
Runs Azure, distributes AI through Microsoft 365 and Copilot, partners commercially with OpenAI, and designs its own Maia accelerators.
Sells business process automation and decisioning software to large organisations.
Tools aimed at particular kinds of work, where proximity to the task matters more than general capability. Developers were the first large professional group to adopt AI tooling at scale.
Provides creative and document software with generative features built into its established tools.
Hosts code and ships Copilot, the coding assistant most developers encountered first.
A biotechnology company applying computational methods within immunotherapy development.
Perception and control models running on hardware in the real world rather than in a data centre. Tesla is the most visible listed example, through driver assistance and autonomy, a planned robotaxi service, the Optimus humanoid programme and custom inference silicon in its vehicles.
Builds electric vehicles and develops driver assistance and autonomy software, a planned robotaxi service, the Optimus humanoid robot programme and custom inference silicon for its vehicles.
Security and governance are not the ninth and tenth layers of a sequence. Every layer adds attack surface, and every layer operates under rules about data, models and trade. They wrap the whole city.
Security is not a step in the sequence. Every layer adds attack surface, from firmware in the data centre to an agent holding credentials on someone's behalf.
Explore →The RulesRules shape what can be built, where data may sit, and what a deployed model must be able to explain about itself. They cut across every layer rather than sitting on top of one.
Explore →The PeopleTechnology creates economic value only when people and organisations use it. Demand runs across every layer rather than sitting on top of one, which is why it is treated here as a cross-cutting question.
Explore →Early ideas that could matter a great deal later, kept deliberately separate from the seven core layers so that speculation is never mistaken for the working ecosystem.
Cheap launch capacity and satellite connectivity extend compute and data collection to places terrestrial networks do not reach. Starlink is the clearest working example, though SpaceX is private and the AI link is indirect rather than core revenue today.
Proposals to place data centres in orbit, using continuous solar power and radiative cooling instead of grid connections and water. Early, unproven and capital hungry, but it is a serious response to the power and cooling constraints on the ground.
The terms used across the diagrams and the book, in plain English. Each one links back to the layer where it does its work.
Six quick questions drawn from the glossary. Read the definition, pick the matching term.
“Software that uses a model to take actions on your behalf, not just answer questions. It can browse, click, buy and send with real permissions.”
Investor & portfolio thinker
Founder, Beaumont Sycamore Media
I spend a lot of my time thinking about where capital flows in technology, both as a practitioner and as an investor trying to understand which businesses will compound over the next decade.
The AI ecosystem is where I spend a lot of my thinking time, both as a practitioner building on these technologies and as an investor trying to understand where the real durable value gets created. There's a lot of noise. Most of it is either hype or fear. The actual story is more interesting than either.
My view is that the infrastructure layers (chips, power, data centres, connectivity) are where the most defensible businesses are being built right now. The application layer will produce enormous winners too, but it's harder to pick them early. The picks and shovels tend to win regardless of who discovers the gold.
This explainer is my attempt to map the ecosystem in plain language. Not for analysts. For anyone who wants to understand what's actually being built and why it matters.