Gartner predicts that by 2030, inference costs for LLMs with 1 trillion parameters will be reduced by more than 90%

TL;DR AI
2 min readKey summary
Gartner says inference costs for trillion-parameter LLMs could fall by more than 90% by 2030 versus 2025, driven by advances in semiconductors and model design.
The firm warns that wider use of AI agents could increase token consumption, offsetting enterprise-wide cost savings.
Gartner recommends using smaller models where possible and reserving large models for tasks that truly need them.
