Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models

TL;DR AI
2 min readKey summary
Researchers introduced SplitQ, a low-bit PTQ method for vision-language models that tackles text-image activation imbalance.
The paper pinpoints uneven modality-specific outlier channels as a major source of quantization error.
SplitQ combines channel splitting with adaptive cross-modal calibration to reduce these errors.
It improves accuracy across multiple datasets and settings, including W4A8, W4A4, W3A3, and W3A2.
The method could make large VLMs more practical on resource-limited devices without losing as much accuracy.
