The artificial intelligence industry is shifting its focus to inference, the process of running trained AI models to produce outputs. This segment is emerging as a significant profit center, distinct from AI model training, which analysts categorize as a cost center. Major chipmakers are now competing to optimize chips for low latency, power efficiency, and cost, driving a trend toward specialized silicon alongside general-purpose GPUs.
Inference’s Unique Demands
Inference has different economic and performance requirements than training. While GPUs deliver strong performance for both, their architecture is often optimized for training, which does not always translate to optimal latency or efficiency for pure inference workloads. Purpose-built inference chips, such as ASICs and other accelerators, can offer faster responses, better energy efficiency, and lower total cost of ownership. Karl Freund, founder and principal analyst at Cambrian AI Research, stated that as a profit center, good latency drives more revenue because users seek rapid responses at the lowest possible cost.
Key takeaways
- The AI chip market is reorganising around inference — running trained models — which analysts treat as a profit centre, against training as a cost centre.
- Nvidia signed a $20 billion licensing deal with Groq, AMD acquired Untether AI's engineering team and later MK1, and Intel is reportedly pursuing a $1.6 billion acquisition of SambaNova.
- A November 2025 Futurum Group survey put GPUs at 58% of data centre compute spending in 2025, but projected XPUs to lead growth in 2026 at 22%, ahead of GPUs at 19% and CPUs at 14%.
- AWS's custom Trainium silicon already handles almost 50% of tokens on its Bedrock inference service.
- Power and cooling limits are the practical driver: a retail branch or edge site cannot host a rack of high-power GPUs, which opens the door to specialised accelerators.
Industry Leaders Pivot to Specialized AI
Chipmakers are actively investing in inference capabilities. Nvidia’s $20 billion licensing deal with Groq highlights this strategic shift. AMD acquired Untether AI’s engineering team and later MK1, an inference startup. Intel is reportedly pursuing a $1.6 billion acquisition of SambaNova and enhances its Xeon CPUs with AMX accelerators, also offering Gaudi AI accelerators. According to Matt Kimball, vice president and principal analyst at Moor Insights & Strategy, larger companies are strengthening their inference portfolios through both products and acquiring engineering talent.
Market Diversification and Enterprise Opportunities
While GPUs, led by Nvidia and AMD, continue to dominate large-scale training and inference, surging demand for inference is creating opportunities for alternatives. Mainstream enterprises scaling AI from pilot programs to production face power, cooling, and GPU supply limitations. These constraints make GPU-heavy clusters impractical for many environments, especially at the edge or in smaller company deployments like retail or branch offices. For instance, a retail environment with limited power and cooling cannot support a rack of high-power GPUs, creating a need for more efficient solutions.
According to a November 2025 Futurum Group survey, GPUs accounted for 58% of data center compute spending in 2025. However, XPUs (processors that are neither GPUs nor CPUs, like ASICs and custom accelerators) are projected to lead growth in 2026 at 22%, surpassing GPUs at 19% and CPUs at 14%. Brendan Burke, research director at Futurum Group, noted that as inference workloads outpace training workloads in token output, demand for diverse architectures will increase due to XPUs’ efficiency advantages in specific inference tasks. Hyperscalers like AWS are also developing custom chips like Trainium for price-performance and efficiency, handling almost 50% of tokens on AWS’s Bedrock inference service.
Startups and Future Outlook
The inference market offers significant opportunities for startups across data centers and edge deployments. Companies like Cerebras, Tenstorrent, FuriosaAI, and Rebellions are developing specialized inference processors. Other startups, including SiFive, NeuReality, and d-Matrix, address memory and networking bottlenecks crucial for inference performance.
Analysts anticipate Nvidia will maintain its overall dominance in both training and inference. However, diverse workload requirements will allow specialized solutions to capture market share. Jim McGregor, principal analyst of Tirias Research, expects further consolidation among startups due to rapid technological shifts. While GPUs remain versatile and programmable, the advantages of specialized inference chips—lower costs, reduced power consumption, and strong performance—create substantial market opportunities. Matt Kimball predicts mainstream enterprise adoption in 2026 will unlock significant demand for inference-centric startups.
Frequently asked questions
What is the difference between AI training and inference?
Training is the one-off process of building a model from data. Inference is running that finished model to produce outputs, over and over, every time a user makes a request. Inference is where ongoing revenue and ongoing cost both sit.
Why can't GPUs simply handle inference too?
They can, and they do. But GPU architecture is optimised for training, which does not always translate into the lowest latency or best energy efficiency for inference. Purpose-built chips can deliver faster responses at lower total cost of ownership.
Does this threaten Nvidia's position?
Analysts quoted expect Nvidia to keep its overall lead across both training and inference. The expectation is that specialised players take share in specific workloads rather than displace GPUs, with consolidation likely among the startups.








