A Visual Explainer

The AI
Ecosystem Explained: the AI infrastructure stack in plain English, layer by layer

Who builds what, and why it matters

The AI Stack, Explained Simply

Think of it
like building
a city.

Every city needs foundations, power, roads, buildings, and people using them. AI is the same, just faster and more expensive.

Written by Matthew Bernath

Scroll to explore

The full picture at a glance

Seven core layers. Two frontier theses. One map of where AI's money actually goes.

Read the ecosystem broadly from infrastructure to application. The layers depend on and reinforce one another, but the real architecture is not perfectly linear. Security, governance and demand are not steps in the sequence: they cut across everything, and they get their own treatment below.

How the layers fit together

Read it bottom to top. Each layer only exists because the ones beneath it do. Security and governance are not steps in the sequence: they wrap the whole thing.

A request travels up the stack. Value flows back down as revenue.

Interactive

Follow it through the ecosystem

Pick a path and see which layers a real question, a real dollar, or a real product actually touches.

What happens when you ask an AI a question

  1. ↓ 1You ask a questionThe People

    A prompt typed into an app

  2. ↓ 2ApplicationThe Shops

    Adds context, permissions and tools

  3. ↓ 3ModelThe Brains

    Turns the prompt into tokens and a response

  4. ↓ 4DataThe Libraries

    Retrieves the records the answer depends on

  5. ↓ 5Cloud and data centreThe Land

    Schedules the work onto real machines

  6. ↓ 6NetworkThe Roads

    Moves data between processors and sites

  7. ↓ 7Accelerator and HBMThe Sand

    Does the arithmetic, fed by memory bandwidth

  8. ↓ 8Electricity and coolingThe Power

    Pays for it in joules and removes the heat

Real architecture is not perfectly linear. Steps run in parallel, repeat, and sometimes get skipped entirely by caching.

01
The Sand

Semiconductors & Manufacturing

AI can run on conventional processors, but specialised accelerators make large modern AI workloads dramatically more efficient. This layer designs, manufactures and packages them.

Silicon waferRAW MATERIALA polished disc of ultra pure silicon, sliced from a crystal grown out of refined sand. Every chip on earth starts here.EUV machineASMLASML's extreme ultraviolet lithography machine prints circuit patterns onto the wafer. Around $200 million each, and nobody else can build one.FoundryTSMCTSMC runs the machines, etches the layers and stacks them into working chips. Most advanced AI silicon is manufactured in Taiwan.Finished GPUNVIDIA / AMD DESIGNNVIDIA and AMD design the chip but own no factories. They send the blueprint to TSMC and sell the result to data centres.PRINTS THE PATTERNETCHES AND STACKSDesign and manufacturing are separate businesses.

Tap or hover a box to see what it does

Sand becomes wafers, wafers become chips. ASML builds the printer, TSMC runs the factory, NVIDIA and AMD design what gets printed.

Accelerators and processors

CPUs are optimised for flexible, general purpose and low latency computing. GPUs contain many more, simpler processing units and are especially effective at performing large numbers of similar calculations in parallel, which is what training and inference mostly consist of.

publicNVDA

NVIDIA

What it does

Designs the GPUs and accelerated computing platforms used for much of modern AI training and inference, and sells the networking that ties them together.

Position in the ecosystem
publicAMD

AMD

What it does

Designs CPUs and the Instinct family of AI accelerators, sold as an alternative to NVIDIA at the data centre scale.

Position in the ecosystem
publicINTC

Intel

What it does

Designs and manufactures CPUs, and is attempting to build a contract manufacturing business for other companies' chip designs.

Position in the ecosystem

Manufacturing and lithography

Designers such as NVIDIA and AMD are fabless: they own no factories. The physical work happens at foundries, using lithography systems that print features measured in nanometres, and packaging that binds logic and memory into a single module.

publicTSM

TSMC

What it does

Manufactures chips designed by other companies, including most leading edge AI accelerators, and provides the advanced packaging that binds logic dies to memory.

Position in the ecosystem
publicASML

ASML

What it does

Builds the lithography systems used to print circuit patterns onto silicon wafers, including extreme ultraviolet machines.

Position in the ecosystem

High bandwidth memory

HBM stacks memory dies vertically and places them beside the accelerator, giving far more bandwidth than conventional memory laid out on a board. Since large models must stream billions of parameters for every token they produce, HBM supply has repeatedly gated how many AI systems can actually be built.

public000660.KS

SK hynix

What it does

Manufactures memory, including high bandwidth memory stacks that sit alongside AI accelerators.

Position in the ecosystem
publicMU

Micron

What it does

Manufactures DRAM, NAND and high bandwidth memory used in AI servers.

Position in the ecosystem
public005930.KS

Samsung Electronics

What it does

Manufactures memory including HBM, runs a contract chip manufacturing business, and builds consumer devices.

Position in the ecosystem

Custom silicon

Google's TPU, Amazon's Trainium and Inferentia, Microsoft's Maia and Meta's MTIA exist because the largest buyers would rather own their cost structure than rent it. Designing in-house improves margins on their own workloads and reduces dependence on a single supplier. For NVIDIA this is both an opportunity, since it still sells into those same data centres, and a risk, since the customers designing these parts are also its biggest source of revenue.

publicAVGO

Broadcom

What it does

Co-designs custom AI accelerators for hyperscalers and supplies much of the switching and connectivity silicon inside data centres.

Position in the ecosystem
publicGOOGL

Alphabet (Google)

What it does

Builds Gemini models through Google DeepMind, designs TPU accelerators, operates Google Cloud, and distributes AI through Search, Workspace and Android.

publicAMZN

Amazon

What it does

Runs AWS, designs Trainium and Inferentia accelerators, hosts third party models through Bedrock, and invests in Anthropic.

Position in the ecosystem
publicMETA

Meta

What it does

Develops the Llama model family and deploys AI across its own products, while designing MTIA accelerators for internal workloads.

Position in the ecosystem
publicMRVL

Marvell Technology

What it does

Designs data infrastructure silicon including optical interconnect, custom compute and storage controllers.

Position in the ecosystem
02
The Power

Energy, Grid & Cooling

AI's infrastructure problem is increasingly not just whether GPUs can be bought, but whether they can be powered and cooled. That runs from generation through the grid to the rack.

Nuclear / SMRBASELOADAlways on generation. Small modular reactors are being pitched as dedicated power plants sitting next to data centres.Gas and solarVARIABLECheap and fast to build, but output moves with weather and fuel prices. Useful, not sufficient on its own for a site that never sleeps.Fuel cellsON SITEBloom Energy style fuel cells generate power at the building itself, bypassing grid queues that can take years to clear.The gridTRANSMISSIONTransmission is now the bottleneck. Connecting a new gigawatt scale site to the grid often takes longer than building the site.Data centreGIGAWATT SCALE DEMANDA single large AI campus can draw as much electricity as a small city, constantly, all year.Compute is electricity in disguise

Tap or hover a box to see what it does

Generation feeds the grid, the grid feeds the data centre. On site fuel cells and small reactors sit alongside it for power the grid cannot guarantee.

Generation

Nuclear, gas, renewables, fuel cells and on site distributed generation. Buyers increasingly want power that is both firm and low carbon, which is a harder combination to source than either on its own.

publicCEG

Constellation Energy

What it does

Operates a large fleet of nuclear generation in the United States and sells power under long term contracts.

Position in the ecosystem
publicVST

Vistra

What it does

Runs a mixed generation fleet including nuclear, gas and storage, and sells into wholesale power markets.

Position in the ecosystem
publicNEE

NextEra Energy

What it does

Develops and operates renewable generation and storage at scale alongside a regulated Florida utility.

Position in the ecosystem
publicBE

Bloom Energy

What it does

Makes solid oxide fuel cells that generate electricity on site from natural gas, biogas or hydrogen.

Position in the ecosystem
publicOKLO

Oklo

What it does

Designing small modular fast reactors intended to sell power directly to large users such as data centres.

Position in the ecosystem

Grid infrastructure

Transmission lines, substations, transformers and the connection agreements that determine whether a site can draw the power it has contracted for. This is where most of the waiting happens.

publicGEV

GE Vernova

What it does

Supplies gas turbines, grid equipment and wind generation, and services the installed base.

Position in the ecosystem
publicABBN.SW

ABB

What it does

Supplies electrification and automation equipment including transformers, switchgear and drives.

Position in the ecosystem

Data centre electrical

Uninterruptible power supplies, switchgear, busway and backup systems sit between the substation and the rack. Availability is engineered here rather than assumed, and no design eliminates outage risk entirely.

publicETN

Eaton

What it does

Supplies electrical distribution equipment, switchgear, busway and uninterruptible power systems.

Position in the ecosystem
publicSU.PA

Schneider Electric

What it does

Supplies power distribution, UPS, cooling and data centre management software, often as a packaged design.

Position in the ecosystem

Cooling and heat rejection

Air cooling struggles as rack densities climb, so liquid cooling has moved from specialist to routine at the high end. Whatever the method, the heat still has to be rejected somewhere, which is why water and site selection keep appearing in the conversation.

publicVRT

Vertiv

What it does

Makes data centre power and thermal management equipment, including liquid cooling for high density racks.

Position in the ecosystem
publicMOD

Modine Manufacturing

What it does

Makes thermal management products including data centre cooling and heat rejection systems.

Position in the ecosystem
03
The Land

Data Centres & Cloud Infrastructure

The buildings, campuses and cloud platforms where AI actually runs. Hyperscale cloud, colocation, specialised AI and HPC facilities, and sovereign infrastructure are four different businesses.

THE BUILDINGRACKED GPUS, COOLED CONSTANTLYHyperscalersAWS / AZURE / GOOGLEAmazon, Microsoft and Alphabet own the largest fleets and rent capacity to everyone else, including their own AI divisions.NeocloudsCOREWEAVEBuilt purely for AI workloads rather than general web hosting. Dense GPU clusters, rented by the hour.Converted mining sitesPOWER ALREADY CONTRACTEDFormer crypto miners already had buildings, cooling and cheap power contracts. Renting that to AI companies pays better than mining.

Tap or hover a box to see what it does

A hall of racked GPUs, cooled constantly, resold as cloud capacity by hyperscalers, neoclouds and converted mining sites.

Hyperscale cloud

The largest platforms sell compute, storage and managed AI services globally, organised into regions and availability zones so that a failure in one place does not take the service with it. Their own AI campuses are increasingly described as AI factories: purpose built sites optimised for training and serving models rather than for general computing.

publicAMZN

Amazon

What it does

Runs AWS, designs Trainium and Inferentia accelerators, hosts third party models through Bedrock, and invests in Anthropic.

Position in the ecosystem
publicGOOGL

Alphabet (Google)

What it does

Builds Gemini models through Google DeepMind, designs TPU accelerators, operates Google Cloud, and distributes AI through Search, Workspace and Android.

publicORCL

Oracle

What it does

Runs Oracle Cloud Infrastructure with a strong GPU cluster business, and sells the databases and applications that hold a great deal of enterprise data.

Position in the ecosystem

Colocation and interconnection

Customers place their own equipment in someone else's building and connect to networks, clouds and each other. Power density per rack is the number that decides whether a given facility can host modern AI hardware at all.

publicEQIX

Equinix

What it does

Operates colocation data centres where customers place their own equipment and interconnect with each other.

Position in the ecosystem
publicDLR

Digital Realty

What it does

Develops and leases large scale data centre capacity to hyperscalers and enterprises.

Position in the ecosystem

Specialised AI and HPC clouds

Operators built specifically around dense GPU clusters, high speed interconnect and the workloads that need them. Sovereign AI infrastructure is a related category, driven by rules about where data and models may physically reside.

publicCRWV

CoreWeave

What it does

Operates GPU cloud infrastructure purpose built for AI training and inference rather than general computing.

Position in the ecosystem
publicNBIS

Nebius Group

What it does

Operates AI focused cloud infrastructure in Europe. It emerged from the old Yandex international business after the Russian operations were sold in 2024, with the Dutch parent keeping the global assets.

Position in the ecosystem

Repurposed and converted sites

Bitcoin mining, and historically Ethereum mining, required many of the same ingredients now sought by AI infrastructure developers: large power allocations, suitable sites and industrial cooling. Ethereum itself moved from proof of work to proof of stake in September 2022, so it is no longer mined.

publicCORZ · IREN · APLD · CIFR

Core Scientific, IREN, Applied Digital, Cipher Mining

What it does

Operators that built large powered sites for cryptocurrency mining and now convert or develop capacity for AI and high performance computing tenants.

Position in the ecosystem
04
The Roads

Networking & Interconnect

Training a large model means tens of thousands of processors behaving as one machine. Networking can become the bottleneck even when there is plenty of compute available.

GPUSSwitchARISTASwitches shuttle data between GPUs inside the building. If the switch is slow, thousands of expensive chips sit waiting.OpticsFIBRE TRANSCEIVERSTransceivers turn electrical signals into light. Coherent and Lumentum make the components that push data down the fibre.Second data centreLarge training runs increasingly span multiple buildings, so the link between sites has to behave almost like an internal cable.Thousands of GPUs must behave as one machine.

Tap or hover a box to see what it does

Inside a rack, switches move data between GPUs. Between buildings, light does the work through fibre optics.

GPU to GPU inside a server

The shortest and fastest hop. NVIDIA's NVLink and NVSwitch connect accelerators inside a server so that several GPUs behave as one larger device, with far more bandwidth than a standard network link could carry.

publicNVDA

NVIDIA

What it does

Designs the GPUs and accelerated computing platforms used for much of modern AI training and inference, and sells the networking that ties them together.

Position in the ecosystem

Server to server and rack to rack

Across a rack and a hall, clusters run on InfiniBand or high speed Ethernet. NVIDIA supplies InfiniBand and its Spectrum-X Ethernet platform following the Mellanox acquisition, while Arista and Broadcom drive the open Ethernet alternative and Credo supplies the high speed connectivity between them.

publicNVDA

NVIDIA

What it does

Designs the GPUs and accelerated computing platforms used for much of modern AI training and inference, and sells the networking that ties them together.

Position in the ecosystem
publicANET

Arista Networks

What it does

Builds high speed Ethernet switching and the network operating system used in large data centres and AI clusters.

Position in the ecosystem
publicAVGO

Broadcom

What it does

Co-designs custom AI accelerators for hyperscalers and supplies much of the switching and connectivity silicon inside data centres.

Position in the ecosystem
publicCRDO

Credo Technology

What it does

Makes high speed connectivity products including active electrical cables and retimers used inside racks.

Position in the ecosystem

Data centre to data centre

Beyond the hall, traffic moves over optics. Transceivers, lasers and photonic components carry campus and long haul links, and co-packaged optics is the direction of travel as data rates climb.

publicCOHR

Coherent

What it does

Makes optical transceivers, lasers and photonic components used to move data between racks and buildings.

Position in the ecosystem
publicLITE

Lumentum

What it does

Supplies photonics including lasers and transceivers for data centre interconnect.

Position in the ecosystem
publicMRVL

Marvell Technology

What it does

Designs data infrastructure silicon including optical interconnect, custom compute and storage controllers.

Position in the ecosystem
publicLWLG

Lightwave Logic

What it does

A pre-revenue research company developing polymer based electro-optic materials for high speed optical modulation.

Position in the ecosystem

Region to region and out to users

Between regions, and then to the people and devices actually making requests. Latency at this distance shapes where inference is served from rather than where models are trained.

publicNET

Cloudflare

What it does

Operates a global network providing content delivery, DDoS protection, zero trust services and edge compute.

Position in the ecosystem
05
The Libraries

Data & Data Infrastructure

The city's libraries, records and archives. Models are only one part of the system: enterprises also need data that is accessible, trusted, governed, structured, searchable, permissioned and connected to business context.

Operational databasesRUNS THE BUSINESSThe systems of record: transactions, customers, inventory. AI that acts on stale data makes expensive mistakes, so freshness is a requirement.Warehouse / lakehouseSNOWFLAKE, DATABRICKSWhere analytical data is consolidated, modelled and governed. The permissions model here decides what an AI system is allowed to see.Vector searchFINDS BY MEANINGEmbeddings let you search by meaning rather than exact wording. This is the piece that lets a model retrieve the right document instead of guessing.Retrieval and contextGROUNDED ANSWERSRetrieval augmented generation: the model is handed the relevant records, with permissions applied, so its answer reflects your business rather than the open internet.The modelANSWERS WITH EVIDENCESame model, better answers, because the context came from governed data instead of a guess.The demo works without this. The deployment does not.

Tap or hover a box to see what it does

Models are only half the system. Warehouses, operational databases and vector search decide whether an answer is grounded in your business or in someone else's.

Warehouses and lakehouses

Where analytical data is consolidated, modelled and governed. The lakehouse pattern merged the cheap storage of a data lake with the governance and query behaviour of a warehouse, mostly on open table formats.

publicSNOW

Snowflake

What it does

Runs a cloud data platform for storing, governing and querying enterprise data, with AI features layered on top.

Position in the ecosystem
private

Databricks

What it does

Provides a lakehouse platform combining data engineering, analytics and machine learning on open table formats.

Public market exposure

Privately held. Investors include Microsoft and several large asset managers, so exposure is indirect and partial.

Position in the ecosystem

Operational databases and streaming

The systems that run the business in real time, plus the pipelines that move events between them. Agents acting on yesterday's data cause expensive mistakes, so freshness is an architectural requirement rather than a nicety.

publicMDB

MongoDB

What it does

Provides a document database, delivered mainly as the Atlas managed service, with integrated vector search.

Position in the ecosystem
publicORCL

Oracle

What it does

Runs Oracle Cloud Infrastructure with a strong GPU cluster business, and sells the databases and applications that hold a great deal of enterprise data.

Position in the ecosystem
publicCFLT

Confluent

What it does

Commercialises Apache Kafka as a managed platform for streaming data between systems.

Position in the ecosystem

Search, vectors and knowledge

Retrieval turns a general model into one that knows about your business. Vector search finds records by meaning rather than exact wording, and metadata, lineage and permissions decide what may be returned to whom.

publicESTC

Elastic

What it does

Provides search, observability and security analytics built on Elasticsearch, including vector search.

Position in the ecosystem
publicPLTR

Palantir

What it does

Sells data integration and decision platforms to governments and large enterprises, increasingly packaged around AI driven workflows.

Position in the ecosystem
publicSAP

SAP

What it does

Provides enterprise resource planning software that runs core finance, supply chain and HR processes for many large firms.

Position in the ecosystem
06
The Brains

Models & AI Platforms

The organisations training foundation models and serving them through APIs and products. Several of the most important are private companies, which is inconvenient but does not make them less important.

Training dataTEXT, CODE, IMAGESVast collections of text, code and media. Quality and licensing of this data is now one of the biggest competitive questions in AI.Training runWEEKS OF COMPUTEThousands of GPUs run flat out for weeks, adjusting billions of parameters. A frontier run can cost hundreds of millions of dollars.Model weightsTHE FINISHED BRAINThe output of training: a large file of numbers. Meta publishes Llama's weights openly, OpenAI keeps GPT's private.InferenceANSWERSRunning the finished model to answer a question. Cheap per request, but it happens billions of times a day, so it dominates demand.ONCE, VERY EXPENSIVEBILLIONS OF TIMESTraining builds the model. Inference is what everyone pays for.

Tap or hover a box to see what it does

Data goes in, training adjusts billions of parameters, and the finished model answers questions billions of times a day.

Frontier labs

Organisations training the largest general purpose models. Several are private, so the practical question for investors is usually which listed business carries meaningful exposure to them.

private

OpenAI

What it does

Develops the GPT family of models and ships them through ChatGPT, an API and enterprise products.

Public market exposure

Private. Microsoft is a major shareholder and its principal cloud and commercial partner, which is the most common route to indirect exposure.

Position in the ecosystem
private

Anthropic

What it does

Develops the Claude family of models, sold through an API, consumer apps and cloud marketplaces.

Public market exposure

Private. Amazon and Alphabet are both significant investors and cloud partners, which is the usual indirect route.

Position in the ecosystem
publicGOOGL

Google DeepMind

What it does

Alphabet's AI research organisation, responsible for the Gemini model family and a long research record beyond language models.

Position in the ecosystem
private

xAI

What it does

Develops the Grok model family and operates its own large training cluster.

Public market exposure

Private, with reported links to other Musk controlled businesses. There is no clean listed proxy.

Position in the ecosystem

Open-weight ecosystems

Model families whose weights can be downloaded and run independently. That is not the same as open source: licences often restrict use, and the Open Source Initiative's definition requires more than published weights.

publicMETA

Meta

What it does

Develops the Llama model family and deploys AI across its own products, while designing MTIA accelerators for internal workloads.

Position in the ecosystem
private

Mistral AI

What it does

Develops efficient models, several released with open weights, aimed at European enterprises and sovereign deployments.

Public market exposure

Private, French, with strategic investors including industrial and technology partners. Exposure is largely unavailable to public investors.

Position in the ecosystem

Platforms and distribution

How models actually reach customers. Cloud marketplaces, enterprise agreements and productivity software are the routes through which most organisations first pay for AI.

publicAMZN

Amazon

What it does

Runs AWS, designs Trainium and Inferentia accelerators, hosts third party models through Bedrock, and invests in Anthropic.

Position in the ecosystem
publicGOOGL

Alphabet (Google)

What it does

Builds Gemini models through Google DeepMind, designs TPU accelerators, operates Google Cloud, and distributes AI through Search, Workspace and Android.

publicNVDA

NVIDIA

What it does

Designs the GPUs and accelerated computing platforms used for much of modern AI training and inference, and sells the networking that ties them together.

Position in the ecosystem
07
The Shops

Applications, Agents & AI Software

Companies that use AI to deliver products, workflows and services people actually buy. This is where model capability meets a budget holder with a problem.

Foundation modelTHE ENGINEA general purpose model such as Gemini, Llama or GPT. Powerful, but useless to a buyer until someone wraps it in a product.Tooling, data plumbing, guardrailsConnecting the model to company data, checking its output, logging what it did and stopping it doing anything expensive or illegal.Enterprise workflowSERVICENOWAI embedded inside the software companies already run their operations on. Low drama, high stickiness.AnalyticsPALANTIRTurning messy institutional data into decisions, mostly for governments, defence and very large enterprises.Vertical appsINSURANCE, BIOTECH, CARSNarrow products aimed at one industry, where domain knowledge matters more than model size.WHAT THE CUSTOMER ACTUALLY BUYS

Tap or hover a box to see what it does

The model is the engine. The application is the car: interface, workflow, data and a customer with a budget.

Enterprise platforms and agents

Software that runs institutional processes, with AI increasingly embedded as agents that take actions rather than only answer questions. Buying criteria here are auditability and integration long before they are model benchmarks.

publicPLTR

Palantir

What it does

Sells data integration and decision platforms to governments and large enterprises, increasingly packaged around AI driven workflows.

Position in the ecosystem
publicNOW

ServiceNow

What it does

Provides a workflow platform for IT, HR and customer operations, with AI agents embedded in those workflows.

Position in the ecosystem
publicCRM

Salesforce

What it does

Sells customer relationship management software with an agent platform layered over its data and workflow estate.

Position in the ecosystem
publicPEGA

Pegasystems

What it does

Sells business process automation and decisioning software to large organisations.

Position in the ecosystem

Creative, developer and vertical software

Tools aimed at particular kinds of work, where proximity to the task matters more than general capability. Developers were the first large professional group to adopt AI tooling at scale.

publicADBE

Adobe

What it does

Provides creative and document software with generative features built into its established tools.

Position in the ecosystem
publicMSFT

GitHub (Microsoft)

What it does

Hosts code and ships Copilot, the coding assistant most developers encountered first.

Position in the ecosystem
publicIBRX

ImmunityBio

What it does

A biotechnology company applying computational methods within immunotherapy development.

Position in the ecosystem

Physical AI

Perception and control models running on hardware in the real world rather than in a data centre. Tesla is the most visible listed example, through driver assistance and autonomy, a planned robotaxi service, the Optimus humanoid programme and custom inference silicon in its vehicles.

publicTSLA

Tesla

What it does

Builds electric vehicles and develops driver assistance and autonomy software, a planned robotaxi service, the Optimus humanoid robot programme and custom inference silicon for its vehicles.

Position in the ecosystem
Across every layer

The Guards and
the Rules.

Security and governance are not the ninth and tenth layers of a sequence. Every layer adds attack surface, and every layer operates under rules about data, models and trade. They wrap the whole city.

Outside the core stack

Two frontier theses

Early ideas that could matter a great deal later, kept deliberately separate from the seven core layers so that speculation is never mistaken for the working ecosystem.

Frontier 08

SpaceX, Starlink and the AI edge

Cheap launch capacity and satellite connectivity extend compute and data collection to places terrestrial networks do not reach. Starlink is the clearest working example, though SpaceX is private and the AI link is indirect rather than core revenue today.

Frontier 09

Orbital compute

Proposals to place data centres in orbit, using continuous solar power and radiative cooling instead of grid connections and water. Early, unproven and capital hungry, but it is a serious response to the power and cooling constraints on the ground.

Glossary

The terms used across the diagrams and the book, in plain English. Each one links back to the layer where it does its work.

Agent
Software that uses a model to take actions on your behalf, not just answer questions. It can browse, click, buy and send with real permissions.
ExampleAn agent that books your travel end to end, not one that just suggests flights.
Related termsFoundation modelInference
07 · Applications, Agents & AI Software
ASIC
Application specific integrated circuit: a chip designed for one job and nothing else. Harder to change than a GPU, but cheaper and faster per unit of work once the design is settled.
ExampleMost custom AI accelerators, including Google's TPU, are ASICs.
Related termsCustom siliconGPU
01 · Semiconductors & Manufacturing
Custom silicon
Chips designed in-house by the big buyers rather than purchased from a merchant supplier. The hyperscalers build their own accelerators to own their cost structure and reduce dependence on one supplier.
ExampleGoogle's TPU, Amazon's Trainium, Microsoft's Maia and Meta's MTIA.
Related termsASICGPUHigh bandwidth memory (HBM)
01 · Semiconductors & Manufacturing
Data centre
A building full of racked computers with industrial power and cooling. AI data centres are built around dense GPU clusters rather than ordinary web servers.
ExampleA single large AI campus can draw as much electricity as a small city.
Related termsGPUHyperscalerThe gridNetwork switch
03 · Data Centres & Cloud Infrastructure
EUV lithography
Extreme ultraviolet lithography, the process that prints circuit patterns onto silicon wafers using light at 13.5 nanometres, just above the X ray range. Only ASML builds the machines.
ExampleOne ASML EUV machine costs around $200 million and ships in 40 freight containers.
Related termsSilicon waferFoundry
01 · Semiconductors & Manufacturing
Fibre optics
Glass threads that carry data as pulses of light. Inside AI data centres and between them, fibre is what keeps thousands of GPUs in step.
ExampleCoherent and Lumentum make the transceivers that turn electricity into light.
Related termsNetwork switchData centre
04 · Networking & Interconnect
Foundation model
A large general purpose model that many products are built on top of. Powerful, but useless to a buyer until someone wraps it in a product.
ExampleGemini, Llama and GPT are foundation models.
Related termsTrainingInferenceModel weights
06 · Models & AI Platforms
Foundry
A factory that manufactures chips designed by someone else. Designers like NVIDIA own no factories; foundries like TSMC own no chip designs.
ExampleTSMC manufactures most of the world's advanced AI silicon in Taiwan.
Related termsSilicon waferEUV lithographyGPU
01 · Semiconductors & Manufacturing
Fuel cell
A device that generates electricity from fuel through chemistry rather than combustion. Used on site at data centres to bypass years long grid connection queues.
ExampleBloom Energy installs fuel cells at the building itself.
Related termsThe gridSmall modular reactor (SMR)
02 · Energy, Grid & Cooling
GPU
Graphics processing unit. Originally built to render games, it turns out the same maths renders neural networks. The workhorse chip of the AI era.
ExampleNVIDIA's data centre GPUs are the default hardware for training frontier models.
Related termsData centreTrainingInference
01 · Semiconductors & Manufacturing
High bandwidth memory (HBM)
Memory dies stacked vertically and placed right beside the accelerator, giving far more bandwidth than conventional memory on a board. Large models must stream billions of parameters for every token, so HBM supply often limits how many AI systems can be built.
ExampleSK hynix, Micron and Samsung are the three main HBM suppliers.
Related termsGPUSilicon wafer
01 · Semiconductors & Manufacturing
Hyperscaler
The handful of companies operating cloud infrastructure at global scale: Amazon, Microsoft and Alphabet. They own the largest GPU fleets and rent capacity to everyone else.
ExampleAWS, Azure and Google Cloud.
Related termsNeocloudData centre
03 · Data Centres & Cloud Infrastructure
Inference
Running a finished model to produce an answer. Cheap per request, but it happens billions of times a day, so it dominates real world compute demand.
ExampleEvery ChatGPT reply is an inference call.
Related termsTrainingGPUFoundation model
06 · Models & AI Platforms
InfiniBand
A very low latency networking standard used to link accelerators inside AI clusters, supplied largely by NVIDIA since it acquired Mellanox. High speed Ethernet is the competing approach.
ExampleLarge training clusters connect thousands of GPUs over InfiniBand.
Related termsNetwork switchFibre opticsGPU
04 · Networking & Interconnect
Liquid cooling
Removing heat by circulating fluid directly to or through the hardware rather than moving air. Once a rack draws enough power, air simply cannot carry the heat away fast enough.
ExampleVertiv and Modine sell liquid cooling systems for high density AI racks.
Related termsData centreThe grid
02 · Energy, Grid & Cooling
Model weights
The output of training: a large file of numbers that is the model. Who owns and shares weights is a strategic choice.
ExampleMeta publishes Llama's weights openly; OpenAI keeps GPT's private.
Related termsFoundation modelTraining
06 · Models & AI Platforms
Neocloud
A cloud provider built purely for AI workloads rather than general web hosting. Dense GPU clusters, rented by the hour.
ExampleCoreWeave.
Related termsHyperscalerGPUData centre
03 · Data Centres & Cloud Infrastructure
Network switch
The hardware that shuttles data between machines. If the switch is slow, thousands of expensive GPUs sit idle waiting for data.
ExampleArista builds the switches inside many AI data centres.
Related termsFibre opticsGPUData centre
04 · Networking & Interconnect
Open weights
A model whose weights can be downloaded and run independently. Not the same as open source: licences often restrict use, and the Open Source Initiative's definition requires more than published weights.
ExampleMeta's Llama family and Mistral's models are distributed this way.
Related termsModel weightsFoundation model
06 · Models & AI Platforms
Retrieval augmented generation (RAG)
Handing the model the relevant records, with permissions applied, before it answers. This is what turns a general model into one that knows about your business.
ExampleA support assistant that reads your policy documents at answer time rather than relying on memory.
Related termsVector searchInference
05 · Data & Data Infrastructure
Silicon wafer
A polished disc of ultra pure silicon, sliced from a crystal grown out of refined sand. Every chip on earth starts as one of these.
ExampleTSMC etches billions of transistors onto each wafer.
Related termsFoundryEUV lithography
01 · Semiconductors & Manufacturing
Small modular reactor (SMR)
A compact nuclear reactor designed to be factory built and installed near the demand. Pitched as dedicated power plants sitting next to data centres.
ExampleOklo is developing small reactors aimed at data centre power.
Related termsThe gridFuel cell
02 · Energy, Grid & Cooling
Sovereign AI
Compute, models and data kept inside a jurisdiction because governments treat them as strategic assets. A large part of why regional AI clouds exist.
ExampleNational AI infrastructure programmes in Europe and the Middle East.
Related termsData centreFoundation model
G2 · Governance, Regulation & Sovereign AI
The grid
The transmission network that moves electricity from power plants to users. For AI, connecting a new gigawatt scale site to the grid often takes longer than building the site.
ExampleUtilities quote multi year waits for new large connections.
Related termsFuel cellSmall modular reactor (SMR)Data centre
02 · Energy, Grid & Cooling
TPU
Tensor processing unit: Google's custom AI accelerator, designed in-house and manufactured by foundries. The longest running proof that a hyperscaler can build credible silicon of its own.
ExampleGoogle trains and serves Gemini on its own TPUs.
Related termsCustom siliconASIC
01 · Semiconductors & Manufacturing
Training
Building a model. Pre-training is the long, expensive phase where thousands of GPUs run flat out for weeks adjusting billions of parameters. Post-training, including fine tuning and alignment, continues after that and is repeated far more often than people assume, so training is an ongoing programme rather than a single event.
ExampleA frontier training run can cost hundreds of millions of dollars.
Related termsInferenceModel weightsGPU
06 · Models & AI Platforms

Test yourself

Six quick questions drawn from the glossary. Read the definition, pick the matching term.

Question 1 of 6Score 0

Software that uses a model to take actions on your behalf, not just answer questions. It can browse, click, buy and send with real permissions.

Frequently Asked

Questions people
actually ask.

No, and this surprises a lot of people. NVIDIA designs the chips and outsources manufacturing, mostly to TSMC in Taiwan. It is a fabless company: engineers and intellectual property rather than factories. AMD and Apple work the same way. Physical manufacturing is extraordinarily capital intensive, and TSMC has spent decades becoming very difficult to replace at the leading edge.Explore this layer: Semiconductors & Manufacturing

CPUs are optimised for flexible, general purpose and low latency computing. They handle branching, unpredictable work and the everyday business of running an operating system. GPUs contain many more, simpler processing units and are especially effective at performing large numbers of similar calculations in parallel. That parallel arithmetic is most of what training and inference consist of, which is why a design originally refined for rendering graphics turned out to suit AI so well.Explore this layer: Semiconductors & Manufacturing

Yes. AI can run on conventional processors, but specialised accelerators make large modern AI workloads dramatically more efficient. The difference is not capability so much as cost, speed and energy. A model that takes seconds on an accelerator might take minutes or hours on a general purpose CPU, and at scale that gap becomes the entire economics of the business.Explore this layer: Semiconductors & Manufacturing

High Bandwidth Memory stacks memory dies vertically and places them next to the accelerator rather than out on the board. That gives far more bandwidth, which matters because a large model has to stream billions of parameters for every token it produces. If memory cannot supply data fast enough, the accelerator waits. That is why memory bandwidth, not raw arithmetic, is often the real constraint, and why HBM supply from SK hynix, Micron and Samsung has repeatedly limited how many AI systems could be built in a given year.Explore this layer: Semiconductors & Manufacturing

Training a large model means running many thousands of accelerators at high utilisation for weeks. Those chips turn nearly all that electricity into heat, and the heat has to be removed, which is where cooling and in some designs water consumption come in. This is why energy, grid connections and cooling have become genuine parts of the AI investment conversation rather than background details. In several markets, securing power is now harder than securing chips.Explore this layer: Energy, Grid & Cooling

Not in the strict sense. Meta's Llama is an open-weight model family. Its weights can be downloaded and run independently, although Meta's licensing terms mean it is not considered open source under the Open Source Definition maintained by the Open Source Initiative. The distinction matters commercially. Open weights let a company run a capable model on its own infrastructure, which puts real pricing pressure on paid APIs, but the licence still places conditions on how it may be used.Explore this layer: Models & AI Platforms

Most AI use is still reactive: you ask, it answers. An agent is given a goal instead, and works out the intermediate steps, calling tools, reading data and looping until it decides the task is finished. That shift is why permissions, audit trails and identity suddenly matter so much. A chatbot that is wrong wastes your time. An agent that is wrong can take an action on your behalf, which is a different category of problem.Explore this layer: Applications, Agents & AI Software

Training is not a single event. Pre-training is the long, expensive phase where a model learns general patterns from very large amounts of data. Post-training shapes that raw model into something useful and better behaved, often using human feedback and reinforcement learning. Fine-tuning adapts an existing model to a narrower task or domain, and can be done repeatedly and comparatively cheaply. Inference is the model actually answering, which happens billions of times a day. Much of the current demand for chips, power and data centres is driven by inference at scale rather than by headline training runs alone.Explore this layer: Models & AI Platforms

Infrastructure overlap. Bitcoin mining, and historically Ethereum mining, required many of the same ingredients now sought by AI infrastructure developers: large power allocations, suitable sites and industrial cooling. Ethereum itself moved from proof of work to proof of stake in September 2022, so it is no longer mined. When mining economics weakened, several operators held energised sites that would take years to permit from scratch. Converting them for AI tenants is a substantial rebuild rather than a relabel, but the scarce ingredient, power, was already secured.Explore this layer: Data Centres & Cloud Infrastructure

Parts of it may be. Some valuations clearly assume everything goes right. But the infrastructure spending is happening regardless of which application company wins, which is the argument for owning the suppliers rather than guessing the winners. The honest position is that the ecosystem contains both durable businesses and speculative ones, and the market is not currently doing a careful job of distinguishing them. That is a reason to read the economics of each layer rather than to treat AI as one trade.Explore this layer: Users, Enterprises & Distribution

A large, general purpose model trained on a very broad dataset, which can then be adapted to particular tasks. The GPT, Claude, Gemini and Llama families are all foundation models. They are called foundations because applications, tools and agents get built on top of them. Training one is an expensive bet that sitting at the base of the stack is where lasting value accrues, which is precisely the question the rest of this site is trying to help you think about.Explore this layer: Models & AI Platforms

Because TSMC, which manufactures most leading edge chips, is based there, and Taiwan sits in a contested region. A serious disruption to its operations would slow the global AI build out considerably, and there is no quick substitute for that capacity. That is the reason the United States, the European Union and Japan are all subsidising domestic manufacturing. The motivation is industrial policy and national security as much as economics.Explore this layer: Semiconductors & Manufacturing

ASML is currently the world's only supplier of production EUV lithography systems, which are required for manufacturing many of the most advanced chips. Other companies make lithography equipment for less advanced processes, so this is not a monopoly on lithography as a whole. The EUV machines themselves are remarkable. Light at a wavelength of 13.5 nanometres, close to the X ray range, is produced by firing a laser at tin droplets tens of thousands of times a second. The light is then steered by mirrors so smooth that if you scaled one to the size of Germany, the biggest bump would be less than a tenth of a millimetre high. A single system costs in the region of two hundred million dollars and ships in dozens of freight containers. Canon and Nikon have both competed in lithography and neither has brought a production EUV system to market. The advantage is technical and cumulative rather than granted by regulators. The Dutch government, under pressure from the United States, also restricts sales of the most advanced systems to China, which makes one company in one Dutch city a genuine chokepoint in the technology relationship between two superpowers.Explore this layer: Semiconductors & Manufacturing
About the Author

Matthew
Bernath

Investor & portfolio thinker
Founder, Beaumont Sycamore Media

I spend a lot of my time thinking about where capital flows in technology, both as a practitioner and as an investor trying to understand which businesses will compound over the next decade.

The AI ecosystem is where I spend a lot of my thinking time, both as a practitioner building on these technologies and as an investor trying to understand where the real durable value gets created. There's a lot of noise. Most of it is either hype or fear. The actual story is more interesting than either.

My view is that the infrastructure layers (chips, power, data centres, connectivity) are where the most defensible businesses are being built right now. The application layer will produce enormous winners too, but it's harder to pick them early. The picks and shovels tend to win regardless of who discovers the gold.

This explainer is my attempt to map the ecosystem in plain language. Not for analysts. For anyone who wants to understand what's actually being built and why it matters.

Read the methodology →

Cover of The AI Ecosystem Explained, second edition, a plain-English guide by Matthew Bernath mapping the layers of the AI infrastructure stack

Premium edition

The book, $19

The AI ecosystem mapped. Every layer and the companies that matter at each one. Learn why and how they got there in the interconnected AI world. Plain English and no jargon. Delivered as a downloadable PDF the moment your payment goes through.

Loading checkout...

Secure checkout via PayPal. You can pay with PayPal or any major card.

Someone makes the chips. Someone powers them. Someone houses them. Someone connects them. Someone organises the data. Someone builds the models. Someone builds on them. And someone has to pay for all of it.