Switch language한국어
Back to the list

NVIDIA AI Releases Nemotron-Labs-Diffusion: A Tri-Mode Language Model with 6× Tokens Per Forward Over Qwen3-8B

TL;DR AI

Key summary

2 min read
  1. NVIDIA released Nemotron-Labs-Diffusion, a 3B/8B/14B model family with base, instruct, and vision-language variants.

  2. The system combines autoregressive decoding, diffusion-based block decoding, and self-speculation under one shared architecture.

  3. It is trained with a joint AR-plus-diffusion objective to balance fast parallel generation with strong language-model accuracy.

  4. The goal is a single model stack that can serve both edge and cloud workloads without changing weights.

Read the original