Switch language한국어
Back to the list

Towards Fine-Grained Robustness: Attention-Guided Test-Time Prompt Tuning for Vision-Language Models

TL;DR AI

Key summary

2 min read
  1. Researchers introduced A-TPT, an attention-guided test-time prompt tuning method for vision-language models like CLIP.

  2. A-TPT refines gradient attention rollout to locate semantically meaningful regions, then uses them to guide spatial augmentations and multi-view inference.

  3. The method improves robustness on both adversarial and clean data, outperforming prior test-time adaptation approaches.

  4. It is especially useful for fine-grained recognition, where preserving semantic information matters for accurate adaptation.

Read the original