VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression
TL;DR AI
2 min readKey summary
VisCo is a parameter-sharing self-compression framework that reuses a pretrained vision-language model as its own compressor.
It encodes visual information into a small set of memory tokens and transfers hierarchical information during decoding.
The method outperforms prior compression approaches across tested ratios and remains stable even at extreme compression levels.
It can also complement the original visual tokens, improving the base model while reducing inference cost and memory overhead.
