Models & AI Platforms
Layer 06 of seven · Models & AI Platforms
- Why it matters
- Model capability sets the ceiling on what every application above can promise. It is also the layer where the cost of competing has risen fastest, which concentrates the field.
- The bottleneck
- Compute, capital and research talent, roughly in that order. Access to large training clusters is a strategic asset in itself.
- Who captures value?
- Labs with distribution, whether their own product surface or a partner's. Selling capability through an API alone is a harder business than selling a product people open every morning.
- What could change?
- Capable open-weight models keep compressing prices from below. If the difference between the best model and a free one narrows for common tasks, value moves upward to applications and data.
The organisations training foundation models and serving them through APIs and products. Several of the most important are private companies, which is inconvenient but does not make them less important.
This is the layer most people mean when they say AI company, and it is the most capital hungry place in the ecosystem to compete. Training a frontier model requires a large cluster, a large research team and a tolerance for spending before knowing whether the result will be best in class for more than a few months. Model version numbers age badly, so it is more useful to think in terms of families and the organisations behind them than in terms of whichever release happens to be current.
Tap or hover a box to see what it does
Frontier labs
Organisations training the largest general purpose models. Several are private, so the practical question for investors is usually which listed business carries meaningful exposure to them.
OpenAI
Develops the GPT family of models and ships them through ChatGPT, an API and enterprise products.
It set the consumer expectation for what an AI assistant is, and consumer familiarity has turned out to be a genuine distribution advantage in enterprise sales.
Extremely high compute costs, capable competitors at lower prices, governance complexity, and regulatory attention.
Private. Microsoft is a major shareholder and its principal cloud and commercial partner, which is the most common route to indirect exposure.
Anthropic
Develops the Claude family of models, sold through an API, consumer apps and cloud marketplaces.
Strong traction with developers and in enterprise settings where reliability and safety posture are part of the buying criteria.
Same capital intensity as every frontier lab, with the added dependence on partners who also compete.
Private. Amazon and Alphabet are both significant investors and cloud partners, which is the usual indirect route.
Google DeepMind
Alphabet's AI research organisation, responsible for the Gemini model family and a long research record beyond language models.
Research depth combined with in-house accelerators and distribution through Google's consumer products is a rare combination.
Turning research advantage into product advantage has not always been straightforward inside a very large company.
xAI
Develops the Grok model family and operates its own large training cluster.
Notable mainly for how quickly it assembled training capacity, which showed that compute build speed is itself a competitive variable.
Funding dependent, closely tied to one founder, and competing against much larger balance sheets.
Private, with reported links to other Musk controlled businesses. There is no clean listed proxy.
Open-weight ecosystems
Model families whose weights can be downloaded and run independently. That is not the same as open source: licences often restrict use, and the Open Source Initiative's definition requires more than published weights.
Meta
Develops the Llama model family and deploys AI across its own products, while designing MTIA accelerators for internal workloads.
Meta's Llama is an open-weight model family. Its weights can be downloaded and run independently, although Meta's licensing terms mean it is not considered open source under the strict Open Source Definition. Releasing capable weights commoditises a layer that rivals sell.
Very large capital spending with returns that arrive through advertising rather than direct AI revenue, and licensing scrutiny.
Mistral AI
Develops efficient models, several released with open weights, aimed at European enterprises and sovereign deployments.
The clearest European answer to the question of who trains models for customers that would rather not send data across the Atlantic.
Competing on efficiency against far larger budgets, and a sovereignty tailwind that is partly political.
Private, French, with strategic investors including industrial and technology partners. Exposure is largely unavailable to public investors.
Platforms and distribution
How models actually reach customers. Cloud marketplaces, enterprise agreements and productivity software are the routes through which most organisations first pay for AI.
Microsoft
Runs Azure, distributes AI through Microsoft 365 and Copilot, partners commercially with OpenAI, and designs its own Maia accelerators.
Few businesses touch as many parts of the ecosystem at once. It owns infrastructure, a commercial route to frontier models, and the enterprise distribution to sell the result.
Very heavy capital spending against uncertain payback, dependence on partner model roadmaps, and buyers questioning per-seat AI pricing.
Amazon
Runs AWS, designs Trainium and Inferentia accelerators, hosts third party models through Bedrock, and invests in Anthropic.
The largest cloud business is trying to own its own accelerator roadmap rather than rent it, which is the clearest example of the custom silicon threat to merchant GPUs.
Custom silicon needs software adoption to matter, and cloud growth is now compared against very large numbers.
Alphabet (Google)
Builds Gemini models through Google DeepMind, designs TPU accelerators, operates Google Cloud, and distributes AI through Search, Workspace and Android.
It is one of the few organisations that designs the chip, runs the data centre, trains the model and owns the consumer surface it ships on.
AI answers can cannibalise the advertising business that funds everything else, and antitrust proceedings could reshape distribution.
NVIDIA
Designs the GPUs and accelerated computing platforms used for much of modern AI training and inference, and sells the networking that ties them together.
The hardware is only half of it. CUDA, the software libraries built on top of it, and a very large developer base make NVIDIA systems the default choice for teams who want to ship rather than port.
Custom accelerators built by its own largest customers, a stronger AMD, revenue concentrated in a handful of buyers, export restrictions, and model architectures that shift the balance of demand.
How to think about the economics
- GrowthVery high
How quickly demand in this part of the ecosystem is expanding.
- Capital intensityVery high
How much money has to be spent up front before revenue arrives.
- Competitive moatModerate
How difficult it is for a credible new entrant to take the business.
- Customer concentrationModerate
How much revenue depends on a small number of buyers.
- Disruption riskVery high
How exposed the layer is to a technical or commercial shift.
This is a framework for thinking about the economics of a layer, not a recommendation. An important AI company, a strategically advantaged company, an investable security and an attractively valued security are four different things.
Where value may accrue
Scale, distribution and switching costs at the platform level. Raw capability leads are usually measured in months, so durable advantage tends to come from where the model is embedded rather than from the model itself.
Key terms in this layer
- Foundation model
- A large general purpose model that many products are built on top of. Powerful, but useless to a buyer until someone wraps it in a product.
- ExampleGemini, Llama and GPT are foundation models.
- Related termsTrainingInferenceModel weights
- Inference
- Running a finished model to produce an answer. Cheap per request, but it happens billions of times a day, so it dominates real world compute demand.
- ExampleEvery ChatGPT reply is an inference call.
- Related termsTrainingGPUFoundation model
- Model weights
- The output of training: a large file of numbers that is the model. Who owns and shares weights is a strategic choice.
- ExampleMeta publishes Llama's weights openly; OpenAI keeps GPT's private.
- Related termsFoundation modelTraining
- Training
- Building a model. Pre-training is the long, expensive phase where thousands of GPUs run flat out for weeks adjusting billions of parameters. Post-training, including fine tuning and alignment, continues after that and is repeated far more often than people assume, so training is an ongoing programme rather than a single event.
- ExampleA frontier training run can cost hundreds of millions of dollars.
- Related termsInferenceModel weightsGPU
- Open weights
- A model whose weights can be downloaded and run independently. Not the same as open source: licences often restrict use, and the Open Source Initiative's definition requires more than published weights.
- ExampleMeta's Llama family and Mistral's models are distributed this way.
- Related termsModel weightsFoundation model
Related questions
Is Meta's Llama really open source?
Not in the strict sense. Meta's Llama is an open-weight model family. Its weights can be downloaded and run independently, although Meta's licensing terms mean it is not considered open source under the Open Source Definition maintained by the Open Source Initiative. The distinction matters commercially. Open weights let a company run a capable model on its own infrastructure, which puts real pricing pressure on paid APIs, but the licence still places conditions on how it may be used.
What is the difference between training and inference?
Training is not a single event. Pre-training is the long, expensive phase where a model learns general patterns from very large amounts of data. Post-training shapes that raw model into something useful and better behaved, often using human feedback and reinforcement learning. Fine-tuning adapts an existing model to a narrower task or domain, and can be done repeatedly and comparatively cheaply. Inference is the model actually answering, which happens billions of times a day. Much of the current demand for chips, power and data centres is driven by inference at scale rather than by headline training runs alone.
What is a foundation model?
A large, general purpose model trained on a very broad dataset, which can then be adapted to particular tasks. The GPT, Claude, Gemini and Llama families are all foundation models. They are called foundations because applications, tools and agents get built on top of them. Training one is an expensive bet that sitting at the base of the stack is where lasting value accrues, which is precisely the question the rest of this site is trying to help you think about.
Doesn't NVIDIA manufacture its own chips?
No, and this surprises a lot of people. NVIDIA designs the chips and outsources manufacturing, mostly to TSMC in Taiwan. It is a fabless company: engineers and intellectual property rather than factories. AMD and Apple work the same way. Physical manufacturing is extraordinarily capital intensive, and TSMC has spent decades becoming very difficult to replace at the leading edge.
What is the difference between a CPU and a GPU?
CPUs are optimised for flexible, general purpose and low latency computing. They handle branching, unpredictable work and the everyday business of running an operating system. GPUs contain many more, simpler processing units and are especially effective at performing large numbers of similar calculations in parallel. That parallel arithmetic is most of what training and inference consist of, which is why a design originally refined for rendering graphics turned out to suit AI so well.
What is agentic AI?
Most AI use is still reactive: you ask, it answers. An agent is given a goal instead, and works out the intermediate steps, calling tools, reading data and looping until it decides the task is finished. That shift is why permissions, audit trails and identity suddenly matter so much. A chatbot that is wrong wastes your time. An agent that is wrong can take an action on your behalf, which is a different category of problem.
Sources and further reading (5)+
- 1Model and platform documentation · OpenAI
- 2Model documentation · Anthropic
- 3Llama community licence agreement · Meta
- 4The Open Source Definition · Open Source Initiative
- 5Research and model announcements · Google DeepMind
The AI ecosystem changes rapidly. Company positions, technologies and market data reflect information available at the date above.