Stability AI Releases Stable Audio 3: A Family of Fast Latent Diffusion Models for Audio Generation and Editing

TL;DR AI
2 min readKey summary
Stability AI released open weights and a paper for Stable Audio 3, a family of latent diffusion audio models.
The models generate 44.1 kHz stereo audio with variable lengths, text conditioning, and inpainting-based editing.
Small and medium weights are available on Hugging Face, while the large model is enterprise-only.
The release brings higher-quality open audio generation and editing, with smaller variants usable on consumer hardware.
