Towards Fine-Grained Robustness: Attention-Guided Test-Time Prompt Tuning for Vision-Language Models

TL;DR AI
2 min readKey summary
Researchers introduced A-TPT, an attention-guided test-time prompt tuning method for vision-language models like CLIP.
A-TPT refines gradient attention rollout to locate semantically meaningful regions, then uses them to guide spatial augmentations and multi-view inference.
The method improves robustness on both adversarial and clean data, outperforming prior test-time adaptation approaches.
It is especially useful for fine-grained recognition, where preserving semantic information matters for accurate adaptation.
