Switch language한국어
Back to the list

[CLS] Is Not Enough: Multi-Label Recognition via Patch-Level Inference and Adaptive Aggregation

TL;DR AI

Key summary

2 min read
  1. Researchers introduced PIAA, a training-free patch-level framework for multi-label image recognition.

  2. PIAA moves beyond the single global [CLS] token by inferring on image patches and adaptively aggregating scores.

  3. The method improves patch discrimination and reduces the vision-language gap with minimal extra compute.

  4. It delivers strong benchmark gains, including more than a 6% mAP boost on NUS-WIDE.

Read the original