Groq’s Inference Chips Are Beating NVIDIA’s Blackwell by 5x on Cost – And Doing It Twice as Fast

TL;DR AI
2 min readKey summary
A Nebius expert told AlphaSense that enterprise AI demand is now driven mostly by inference, not training.
Buyers are increasingly comparing chips by cost per million tokens instead of GPU-hour pricing.
In that model, Groq is estimated at about $0.05-$0.10 per million tokens and around 800 tokens per second.
NVIDIA’s Blackwell is estimated at about $0.25 per million tokens and around 450 tokens per second, making Groq look cheaper and faster for high-volume inference.
The shift could boost specialized inference chips and put pressure on NVIDIA’s AI infrastructure dominance.

