Switch language한국어
Back to the list

NVIDIA launches distilled model for 'Cosmos3 4-Step'... "Best open-weight performance even with 90% less inference"

TL;DR AI

Key summary

2 min read
  1. NVIDIA released distilled 4-step open-weight text-to-image and image-to-video models on Hugging Face, with commercial use allowed.

  2. The new models cut inference steps from 50 to 4 for text-to-image and from 35 to 4 for image-to-video, while removing classifier-free guidance to improve efficiency.

  3. In benchmark rankings, the image-to-video model placed first among open-weight models, and the text-to-image model ranked third.

  4. The release highlights how generative AI can reduce latency and compute cost without giving up quality, especially for synthetic data and physical AI workflows.

Read the original