Generative AI is the part of the wider AI ecosystem pulled along by models that create new text, code, images, audio and video. It runs on the same infrastructure as everything else, but the money moves through a distinct value chain. This page traces that chain from accelerators to agents.
Last fact checked 29 August 2026
Five stages carry a generated answer from raw silicon to a person. Each one maps onto a layer of the full stack, so you can read the detail wherever you want more depth.
Generative models are matrix maths at enormous scale. Training a frontier model and then serving it to millions of people both come back to accelerators, high bandwidth memory and the packaging that binds them together.
Someone has to own the buildings, the cooling and the power contracts. Generative AI shifted the shape of that demand from lots of small servers to dense clusters designed around a single training run.
Pre-training corpora, licensed content, vector search and the pipelines that keep a model connected to current company data. Most generative products fail on retrieval quality long before they fail on model quality.
Closed frontier models sold through an API, open-weight models you can host yourself, and the platforms that host, fine-tune and evaluate both. This is the layer people usually mean when they say generative AI.
Assistants, copilots, image and video tools, coding agents and the workflow software quietly adding generation into features people already use. Value here depends on distribution, not just capability.
Plenty of AI infrastructure has nothing to do with generation. Forecasting, fraud scoring, recommendation and industrial vision all sit on the same chips and data centres. Treating the two as identical is the most common mistake in the market, because it makes a specific product wave look like the entire industry.
Security, regulation and buyer demand cut across all of them: Cybersecurity, Governance, Regulation & Sovereign AI, Users, Enterprises & Distribution.
Designs the GPUs and accelerated computing platforms used for much of modern AI training and inference, and sells the networking that ties them together.
Runs Azure, distributes AI through Microsoft 365 and Copilot, partners commercially with OpenAI, and designs its own Maia accelerators.
Runs AWS, designs Trainium and Inferentia accelerators, hosts third party models through Bedrock, and invests in Anthropic.
Builds Gemini models through Google DeepMind, designs TPU accelerators, operates Google Cloud, and distributes AI through Search, Workspace and Android.
Develops the GPT family of models and ships them through ChatGPT, an API and enterprise products.
Develops the Claude family of models, sold through an API, consumer apps and cloud marketplaces.
Runs Azure, distributes AI through Microsoft 365 and Copilot, partners commercially with OpenAI, and designs its own Maia accelerators.
Runs AWS, designs Trainium and Inferentia accelerators, hosts third party models through Bedrock, and invests in Anthropic.
Builds Gemini models through Google DeepMind, designs TPU accelerators, operates Google Cloud, and distributes AI through Search, Workspace and Android.
Develops the GPT family of models and ships them through ChatGPT, an API and enterprise products.
Sells data integration and decision platforms to governments and large enterprises, increasingly packaged around AI driven workflows.
Provides a workflow platform for IT, HR and customer operations, with AI agents embedded in those workflows.
It is the chain of businesses required to turn electricity and silicon into generated text, code, images, audio and video. Compute, capacity, data, models and applications, with security, regulation and buyer demand cutting across all of it.
The infrastructure stack describes everything that makes modern AI work, including forecasting, recommendation and vision workloads. The generative slice is narrower: it is the part of that stack pulled along by foundation models that produce new content rather than only classify or predict.
A large share of spend from generative AI applications flows back down to model providers, cloud capacity and ultimately accelerators and power. Application companies keep the customer relationship, but their cost of goods sits several layers below them.
Training creates the model and is a large, occasional capital-like cost. Inference is running the finished model for every user request, and it is a recurring cost that scales with usage. Products that grow quickly feel inference cost far more than training cost.
Yes. Open-weight models such as Llama shift where value accrues rather than removing the layers beneath. You still need accelerators, hosting, data pipelines and evaluation, you simply run more of it yourself.