NVIDIA launches distilled model for 'Cosmos3 4-Step'... "Best open-weight performance even with 90% less inference"

TL;DR AI
2 min readKey summary
NVIDIA released distilled 4-step open-weight text-to-image and image-to-video models on Hugging Face, with commercial use allowed.
The new models cut inference steps from 50 to 4 for text-to-image and from 35 to 4 for image-to-video, while removing classifier-free guidance to improve efficiency.
In benchmark rankings, the image-to-video model placed first among open-weight models, and the text-to-image model ranked third.
The release highlights how generative AI can reduce latency and compute cost without giving up quality, especially for synthetic data and physical AI workflows.
