כתבה
arXiv cs.LG ·
ברידג' ויז'ן: פריורים של מודלי תצוגה לצד CLIP לזיהוי-עומק של תפנים בתמונות רפואיות
Bridging Vision Foundation Model Priors with CLIP for Spatial-aware Few-shot Anomaly Detection in Medical Images
מודלי Vision-Language מקשרים בין CLIP לזיהוי-עומק של תפנים בתמונות רפואיות באמצעות זיהוי-עומק של תפנים.
תקציר מקורי באנגליתarXiv:2609.12454v1 Announce Type: cross Abstract: Vision-Language Models such as CLIP enable effective few-shot medical anomaly detection (AD) via strong image-text semantic alignment. However, their globally contrastive pretraining lacks explicit spatial supervision, limiting precise lesion localization. In contrast, Vision Foundation Models (VFMs) such as DINO learn spatially coherent patch representations via self-distillation and local-to-global consistency, better capturing fine-grained anatomical structures. Leveraging this complementarity, we propose Spatial-FAD, a spatial-aware few-shot medical AD framework that improves lesion localization by combining VFM spatial priors with CLIP semantics. Specifically, we introduce a VFM-enhanced adapter that injects a structural affinity prior
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית