GOTS: Greedy Orthogonal Token Selection for High-Resolution Vision-Language Models

TL;DR AI
2 min readKey summary
Researchers introduced GOTS, a training-free method for reducing visual tokens in high-resolution vision-language models.
GOTS greedily keeps tokens with the largest orthogonal residual energy relative to the selected set.
It shows strong results across backbones and benchmarks such as Qwen-VL, InternVL, and OCRBench, even after selection overhead is counted.
The approach lowers inference cost and latency while preserving more performance than competing token-pruning methods.
