NVIDIA releases the diffusion language model "Nemotron-Labs-Diffusion," along with a VLM that can process images and switch between diffusion and autoregressive modes

TL;DR AI
2 min readKey summary
NVIDIA released Nemotron-Labs-Diffusion, an open model family with three inference modes: autoregressive, diffusion, and self-speculation.
The 8B model is reported to beat Qwen3-8B on several benchmarks while running much faster in self-speculation mode.
A vision-language version can also process images, broadening the models' multimodal use cases.
The launch signals NVIDIA's push toward hybrid architectures that aim to blend quality and speed in generative AI.



