כתבה
arXiv cs.LG ·
אוגמנטציה של טוקן CLS עם טוקנים ROI למציאת תמונות דקה
Region-Aware CLS Token Augmentation for Fine-Grained Image Retrieval
במאמר זה, המחברים מציגים שיטה לאוגמנטציה של טוקן CLS עם טוקנים ROI למציאת תמונות דקה. השיטה משתמשת ב-DINOv2-reg ומציעה פתרון לבעיית המציאת תמונות דקה.
תקציר מקורי באנגליתarXiv:2610.10991v1 Announce Type: cross Abstract: Image retrieval methods often rely on a single global semantic descriptor extracted from an image, e.g., the [CLS] token in vision transformers. However, trying to squeeze all the semantic information of an image into a single descriptor can hurt downstream retrieval performance, especially for fine-grained retrieval tasks. In this work, we augment the semantic tokens in the newer visual transformers, the global [CLS] token and the four register tokens, with a carefully selected collection of spatial tokens, aiming to capture the spatial region representation that characterizes the contents captured in each of the semantic tokens. We leverage the DINOv2-reg model, which includes register tokens that emergently learn object and part-based re
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית