OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs
TL;DR AI
2 min readKey summary
Researchers introduced OmniDelta, a training-free framework for allocating compression budgets in omni-modal LLMs.
It combines intent-aware inter-modal allocation with content-aware intra-modal allocation to decide how many audio and video tokens to keep.
On four benchmarks with Qwen2.5-Omni, it improved the accuracy-efficiency tradeoff.
At 25% token retention, it reduced GPU memory by 22.0% and sped up inference by 1.64x.
