NVIDIA AI Releases Nemotron-Labs-Diffusion: A Tri-Mode Language Model with 6× Tokens Per Forward Over Qwen3-8B

TL;DR AI
2 min readKey summary
NVIDIA released Nemotron-Labs-Diffusion, a 3B/8B/14B model family with base, instruct, and vision-language variants.
The system combines autoregressive decoding, diffusion-based block decoding, and self-speculation under one shared architecture.
It is trained with a joint AR-plus-diffusion objective to balance fast parallel generation with strong language-model accuracy.
The goal is a single model stack that can serve both edge and cloud workloads without changing weights.
