The Brains

Models & AI Platforms

Layer 06 of seven · Models & AI Platforms

Why it matters
Model capability sets the ceiling on what every application above can promise. It is also the layer where the cost of competing has risen fastest, which concentrates the field.
The bottleneck
Compute, capital and research talent, roughly in that order. Access to large training clusters is a strategic asset in itself.
Who captures value?
Labs with distribution, whether their own product surface or a partner's. Selling capability through an API alone is a harder business than selling a product people open every morning.
What could change?
Capable open-weight models keep compressing prices from below. If the difference between the best model and a free one narrows for common tasks, value moves upward to applications and data.

The organisations training foundation models and serving them through APIs and products. Several of the most important are private companies, which is inconvenient but does not make them less important.

This is the layer most people mean when they say AI company, and it is the most capital hungry place in the ecosystem to compete. Training a frontier model requires a large cluster, a large research team and a tolerance for spending before knowing whether the result will be best in class for more than a few months. Model version numbers age badly, so it is more useful to think in terms of families and the organisations behind them than in terms of whichever release happens to be current.

Training dataTEXT, CODE, IMAGESVast collections of text, code and media. Quality and licensing of this data is now one of the biggest competitive questions in AI.Training runWEEKS OF COMPUTEThousands of GPUs run flat out for weeks, adjusting billions of parameters. A frontier run can cost hundreds of millions of dollars.Model weightsTHE FINISHED BRAINThe output of training: a large file of numbers. Meta publishes Llama's weights openly, OpenAI keeps GPT's private.InferenceANSWERSRunning the finished model to answer a question. Cheap per request, but it happens billions of times a day, so it dominates demand.ONCE, VERY EXPENSIVEBILLIONS OF TIMESTraining builds the model. Inference is what everyone pays for.

Tap or hover a box to see what it does

Data goes in, training adjusts billions of parameters, and the finished model answers questions billions of times a day.

Frontier labs

Organisations training the largest general purpose models. Several are private, so the practical question for investors is usually which listed business carries meaningful exposure to them.

private

OpenAI

What it does

Develops the GPT family of models and ships them through ChatGPT, an API and enterprise products.

Why it matters

It set the consumer expectation for what an AI assistant is, and consumer familiarity has turned out to be a genuine distribution advantage in enterprise sales.

What could go wrong

Extremely high compute costs, capable competitors at lower prices, governance complexity, and regulatory attention.

Public market exposure

Private. Microsoft is a major shareholder and its principal cloud and commercial partner, which is the most common route to indirect exposure.

Position in the ecosystem
private

Anthropic

What it does

Develops the Claude family of models, sold through an API, consumer apps and cloud marketplaces.

Why it matters

Strong traction with developers and in enterprise settings where reliability and safety posture are part of the buying criteria.

What could go wrong

Same capital intensity as every frontier lab, with the added dependence on partners who also compete.

Public market exposure

Private. Amazon and Alphabet are both significant investors and cloud partners, which is the usual indirect route.

Position in the ecosystem
publicGOOGL

Google DeepMind

What it does

Alphabet's AI research organisation, responsible for the Gemini model family and a long research record beyond language models.

Why it matters

Research depth combined with in-house accelerators and distribution through Google's consumer products is a rare combination.

What could go wrong

Turning research advantage into product advantage has not always been straightforward inside a very large company.

Position in the ecosystem
private

xAI

What it does

Develops the Grok model family and operates its own large training cluster.

Why it matters

Notable mainly for how quickly it assembled training capacity, which showed that compute build speed is itself a competitive variable.

What could go wrong

Funding dependent, closely tied to one founder, and competing against much larger balance sheets.

Public market exposure

Private, with reported links to other Musk controlled businesses. There is no clean listed proxy.

Position in the ecosystem

Open-weight ecosystems

Model families whose weights can be downloaded and run independently. That is not the same as open source: licences often restrict use, and the Open Source Initiative's definition requires more than published weights.

publicMETA

Meta

What it does

Develops the Llama model family and deploys AI across its own products, while designing MTIA accelerators for internal workloads.

Why it matters

Meta's Llama is an open-weight model family. Its weights can be downloaded and run independently, although Meta's licensing terms mean it is not considered open source under the strict Open Source Definition. Releasing capable weights commoditises a layer that rivals sell.

What could go wrong

Very large capital spending with returns that arrive through advertising rather than direct AI revenue, and licensing scrutiny.

Position in the ecosystem
private

Mistral AI

What it does

Develops efficient models, several released with open weights, aimed at European enterprises and sovereign deployments.

Why it matters

The clearest European answer to the question of who trains models for customers that would rather not send data across the Atlantic.

What could go wrong

Competing on efficiency against far larger budgets, and a sovereignty tailwind that is partly political.

Public market exposure

Private, French, with strategic investors including industrial and technology partners. Exposure is largely unavailable to public investors.

Position in the ecosystem

Platforms and distribution

How models actually reach customers. Cloud marketplaces, enterprise agreements and productivity software are the routes through which most organisations first pay for AI.

publicMSFT

Microsoft

What it does

Runs Azure, distributes AI through Microsoft 365 and Copilot, partners commercially with OpenAI, and designs its own Maia accelerators.

Why it matters

Few businesses touch as many parts of the ecosystem at once. It owns infrastructure, a commercial route to frontier models, and the enterprise distribution to sell the result.

What could go wrong

Very heavy capital spending against uncertain payback, dependence on partner model roadmaps, and buyers questioning per-seat AI pricing.

publicAMZN

Amazon

What it does

Runs AWS, designs Trainium and Inferentia accelerators, hosts third party models through Bedrock, and invests in Anthropic.

Why it matters

The largest cloud business is trying to own its own accelerator roadmap rather than rent it, which is the clearest example of the custom silicon threat to merchant GPUs.

What could go wrong

Custom silicon needs software adoption to matter, and cloud growth is now compared against very large numbers.

Position in the ecosystem
publicGOOGL

Alphabet (Google)

What it does

Builds Gemini models through Google DeepMind, designs TPU accelerators, operates Google Cloud, and distributes AI through Search, Workspace and Android.

Why it matters

It is one of the few organisations that designs the chip, runs the data centre, trains the model and owns the consumer surface it ships on.

What could go wrong

AI answers can cannibalise the advertising business that funds everything else, and antitrust proceedings could reshape distribution.

publicNVDA

NVIDIA

What it does

Designs the GPUs and accelerated computing platforms used for much of modern AI training and inference, and sells the networking that ties them together.

Why it matters

The hardware is only half of it. CUDA, the software libraries built on top of it, and a very large developer base make NVIDIA systems the default choice for teams who want to ship rather than port.

What could go wrong

Custom accelerators built by its own largest customers, a stronger AMD, revenue concentrated in a handful of buyers, export restrictions, and model architectures that shift the balance of demand.

Position in the ecosystem

How to think about the economics

  • GrowthVery high

    How quickly demand in this part of the ecosystem is expanding.

  • Capital intensityVery high

    How much money has to be spent up front before revenue arrives.

  • Competitive moatModerate

    How difficult it is for a credible new entrant to take the business.

  • Customer concentrationModerate

    How much revenue depends on a small number of buyers.

  • Disruption riskVery high

    How exposed the layer is to a technical or commercial shift.

This is a framework for thinking about the economics of a layer, not a recommendation. An important AI company, a strategically advantaged company, an investable security and an attractively valued security are four different things.

Where value may accrue

Scale, distribution and switching costs at the platform level. Raw capability leads are usually measured in months, so durable advantage tends to come from where the model is embedded rather than from the model itself.

Key terms in this layer

Foundation model
A large general purpose model that many products are built on top of. Powerful, but useless to a buyer until someone wraps it in a product.
ExampleGemini, Llama and GPT are foundation models.
Related termsTrainingInferenceModel weights
Inference
Running a finished model to produce an answer. Cheap per request, but it happens billions of times a day, so it dominates real world compute demand.
ExampleEvery ChatGPT reply is an inference call.
Related termsTrainingGPUFoundation model
Model weights
The output of training: a large file of numbers that is the model. Who owns and shares weights is a strategic choice.
ExampleMeta publishes Llama's weights openly; OpenAI keeps GPT's private.
Related termsFoundation modelTraining
Training
Building a model. Pre-training is the long, expensive phase where thousands of GPUs run flat out for weeks adjusting billions of parameters. Post-training, including fine tuning and alignment, continues after that and is repeated far more often than people assume, so training is an ongoing programme rather than a single event.
ExampleA frontier training run can cost hundreds of millions of dollars.
Related termsInferenceModel weightsGPU
Open weights
A model whose weights can be downloaded and run independently. Not the same as open source: licences often restrict use, and the Open Source Initiative's definition requires more than published weights.
ExampleMeta's Llama family and Mistral's models are distributed this way.
Related termsModel weightsFoundation model

Related questions

Is Meta's Llama really open source?

Not in the strict sense. Meta's Llama is an open-weight model family. Its weights can be downloaded and run independently, although Meta's licensing terms mean it is not considered open source under the Open Source Definition maintained by the Open Source Initiative. The distinction matters commercially. Open weights let a company run a capable model on its own infrastructure, which puts real pricing pressure on paid APIs, but the licence still places conditions on how it may be used.

What is the difference between training and inference?

Training is not a single event. Pre-training is the long, expensive phase where a model learns general patterns from very large amounts of data. Post-training shapes that raw model into something useful and better behaved, often using human feedback and reinforcement learning. Fine-tuning adapts an existing model to a narrower task or domain, and can be done repeatedly and comparatively cheaply. Inference is the model actually answering, which happens billions of times a day. Much of the current demand for chips, power and data centres is driven by inference at scale rather than by headline training runs alone.

What is a foundation model?

A large, general purpose model trained on a very broad dataset, which can then be adapted to particular tasks. The GPT, Claude, Gemini and Llama families are all foundation models. They are called foundations because applications, tools and agents get built on top of them. Training one is an expensive bet that sitting at the base of the stack is where lasting value accrues, which is precisely the question the rest of this site is trying to help you think about.

Doesn't NVIDIA manufacture its own chips?

No, and this surprises a lot of people. NVIDIA designs the chips and outsources manufacturing, mostly to TSMC in Taiwan. It is a fabless company: engineers and intellectual property rather than factories. AMD and Apple work the same way. Physical manufacturing is extraordinarily capital intensive, and TSMC has spent decades becoming very difficult to replace at the leading edge.

What is the difference between a CPU and a GPU?

CPUs are optimised for flexible, general purpose and low latency computing. They handle branching, unpredictable work and the everyday business of running an operating system. GPUs contain many more, simpler processing units and are especially effective at performing large numbers of similar calculations in parallel. That parallel arithmetic is most of what training and inference consist of, which is why a design originally refined for rendering graphics turned out to suit AI so well.

What is agentic AI?

Most AI use is still reactive: you ask, it answers. An agent is given a goal instead, and works out the intermediate steps, calling tools, reading data and looping until it decides the task is finished. That shift is why permissions, audit trails and identity suddenly matter so much. A chatbot that is wrong wastes your time. An agent that is wrong can take an action on your behalf, which is a different category of problem.

Sources and further reading (5)+
  1. 1Model and platform documentation · OpenAI
  2. 2Model documentation · Anthropic
  3. 3Llama community licence agreement · Meta
  4. 4The Open Source Definition · Open Source Initiative
  5. 5Research and model announcements · Google DeepMind
Last fact-checked: 29 August 2026

The AI ecosystem changes rapidly. Company positions, technologies and market data reflect information available at the date above.