Inference, not training, is the cost line
As AI products reach volume, attention is shifting from training runs to the recurring cost of serving. Routing, compression and purpose-built silicon are where margin is won.
The signal
Hardware and serving startups are marketing on cost per token at a quality bar rather than on peak throughput.
Industries involved
Technologies
SemiconductorsScaling
Inference Optimisation
The set of techniques — quantisation, batching, caching, distillation and specialised silicon — that reduce the cost and latency of running a trained model.
Artificial IntelligenceScaling
Foundation Models
Large models trained on broad data that can be adapted to many downstream tasks rather than being built for one purpose.