Switch language한국어
Back to the list

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception

TL;DR AI

Key summary

2 min read
  1. Researchers introduced UniAudio-Token, a new audio tokenizer that extends speech-focused semantic tokenizers to broader audio perception.

  2. It preserves the single-codebook semantic approach but reduces acoustic information loss with Semantic-Acoustic Primitives and a content-aware gating method called Semantic-Acoustic Equilibrium.

  3. Evaluated as a unified audio interface, it outperforms single-codebook baselines on both understanding and generation benchmarks.

  4. The framework aims to improve audio-LLM interfaces by handling more than speech while keeping high-quality speech output.

Read the original