MiniMax teases M3 model with new sparse attention mechanism, 15.6x long-context response speed boost

TL;DR AI
2 min readKey summary
MiniMax published a technical report on its M2 language models, detailing architecture and efficiency techniques.
The company also teased M3, a new model series built around a custom sparse attention framework.
MiniMax says the design can speed up million-token long-context decoding by up to 15.6x.
If validated, the approach could make ultra-long-context AI systems cheaper and faster for enterprise and agent workloads.
