Gartner Predicts 90%+ Inference Cost Cut for 1-Trillion-Parameter LLMs by 2030

TL;DR AI
2 min readKey summary
Gartner predicts that inference costs for 1-trillion-parameter LLMs will drop by over 90% by 2030 versus 2025.
The firm says the reduction will come from combined improvements in semiconductors, infrastructure, model design, chip utilization, and inference silicon.
Gartner presented two scenarios—Frontier (leading-edge chips) and Legacy Blend (existing semiconductor benchmarks)—with Legacy Blend showing higher absolute costs.
Gartner warned cost-per-inference cuts may not lower enterprise AI spending because AI agents could process 5–30x more tokens.
The firm advised using small domain-specific models for frequent tasks and large LLMs only for complex processing.



