Forecast

Gartner: Cloud Spending on Language Models Set to Nearly Double

Gartner, cloud spending, AI-optimized IaaS, AI-optimized IaaS spending forecast 2026, Gartner AI infrastructure forecast 2026, Gartner forecast
Facebook
X
LinkedIn
Reddit
WhatsApp

Gartner expects spending on specialized IaaS services to reach $42 billion in 2026. The main driver is no longer training large language models, but running them in real-world production environments.

The generative AI boom continues to fuel demand for specialized cloud infrastructure. According to a new forecast from analyst firm Gartner, worldwide spending on so-called AI-optimized Infrastructure as a Service (IaaS) is expected to reach $42.3 billion by the end of 2026. That would represent growth of 96.4 percent compared with the previous year, when the market had already expanded by 180 percent to $21.5 billion.

Ad

Gartner defines AI-optimized IaaS as cloud infrastructure specifically designed for training and running AI models. This includes GPU and accelerator instances as well as purpose-built networking and storage architectures. By comparison, the overall IaaS market, which also includes traditional cloud servers without an AI focus, is expected to grow at a much slower rate of 29.3 percent in 2026, reaching approximately $287 billion.

For 2027, Gartner expects AI-optimized infrastructure spending to grow by another 56.5 percent, albeit at a slower pace, reaching $66.1 billion.

Inference Overtakes Training

One of the most notable shifts within the market is the changing balance between training and inference. For the first time, spending on inference, meaning the use of already trained models in production, is expected to surpass spending on training new models. Gartner forecasts $23.3 billion in inference workloads in 2026, compared with $19 billion for training. Inference will therefore account for 55 percent of total AI IaaS spending, rising to 59 percent in 2027.

Ad

Hardeep Singh, Senior Principal Research Analyst at Gartner, attributes the shift to the move from model development to production deployment. Fine-tuned, domain-specific models are increasingly being integrated into customer-facing and operational systems, where they need to run continuously and in real time rather than consuming compute capacity only intermittently during training. According to Singh, this transition is creating “sustained demand for AI-optimized infrastructure.”

Gartner also identifies the rise of agentic AI as another key driver. Autonomous systems capable of handling multistep tasks on their own significantly increase the amount of compute required per request. This makes inference infrastructure a potential bottleneck and a critical building block for enterprises looking to scale their AI strategies across the organization.

The figures reinforce a trend that has been emerging for some time. Cloud providers such as AWS, Microsoft Azure and Google Cloud are investing heavily in specialized hardware to support not only the training of large models, but increasingly their continuous operation in production environments. Whether this pace of growth can be sustained amid rising energy and hardware costs will become clearer over the coming quarters.

(Editorial Team)

Ad

Artikel zu diesem Thema

Weitere Artikel