[CLS] Is Not Enough: Multi-Label Recognition via Patch-Level Inference and Adaptive Aggregation

TL;DR AI
2 min readKey summary
Researchers introduced PIAA, a training-free patch-level framework for multi-label image recognition.
PIAA moves beyond the single global [CLS] token by inferring on image patches and adaptively aggregating scores.
The method improves patch discrimination and reduces the vision-language gap with minimal extra compute.
It delivers strong benchmark gains, including more than a 6% mAP boost on NUS-WIDE.
