UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception

TL;DR AI
2 min readKey summary
Researchers introduced UniAudio-Token, a new audio tokenizer that extends speech-focused semantic tokenizers to broader audio perception.
It preserves the single-codebook semantic approach but reduces acoustic information loss with Semantic-Acoustic Primitives and a content-aware gating method called Semantic-Acoustic Equilibrium.
Evaluated as a unified audio interface, it outperforms single-codebook baselines on both understanding and generation benchmarks.
The framework aims to improve audio-LLM interfaces by handling more than speech while keeping high-quality speech output.
