Switch language한국어
Back to the list

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

TL;DR AI

Key summary

2 min read
  1. Researchers introduced OmniDelta, a training-free framework for allocating compression budgets in omni-modal LLMs.

  2. It combines intent-aware inter-modal allocation with content-aware intra-modal allocation to decide how many audio and video tokens to keep.

  3. On four benchmarks with Qwen2.5-Omni, it improved the accuracy-efficiency tradeoff.

  4. At 25% token retention, it reduced GPU memory by 22.0% and sped up inference by 1.64x.

Read the original