כתבה
arXiv cs.AI ·
דמיון גאומטרי בייצוגי VLM
Geometric Similarity in VLM Low-Level Vision Representations
חוקרים הציגו את מסגרת GeoSim, הבוחנת דמיון גאומטרי בייצוגי VLM. המחקר בודק 24 משימות שונות וחושף עקרונות מארגנים של ייצוגים חזותיים. התוצאות מראות גבולות בהסכמה בין משימות ומודלים.
תקציר מקורי באנגליתarXiv:2610.00848v1 Announce Type: cross Abstract: Vision-language models (VLMs) have emerged as powerful candidates for universal vision backbones, with representative architectures including autoregressive (AR) models and diffusion transformers (DiTs). Yet, adapting them efficiently for all-in-one low-level image restoration remains a challenge. Crucially, the field lacks an understanding of how VLMs organize hidden-layer representations and whether these structurally distinct paradigms share a common geometric organization for pixel-level perception. Such shared organization is a prerequisite for building highly transferable, unified restoration VLMs and adapters. In this paper, we systematically investigate representational similarity across 24 low-level tasks spanning 5 categories. We
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית