Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model

TL;DR AI
2 min readKey summary
Thinking Machines Lab has released Inkling-Small, an Apache 2.0 open-weights multimodal MoE model.
It handles text, images, and audio, with 276B total parameters and a 1M-token context window.
Trained on NVIDIA GB300 NVL72 systems, it can also run when quantized on a single B300 or dual H200 setup.
Benchmark results show gains on several reasoning and coding tasks, though some knowledge and epistemic measures lag.
The release expands practical self-hosted options for startups and enterprises needing private multimodal AI.



