כתבה
arXiv cs.AI ·
חזרה למערת פלאטו: בדיקת התכנסות ייצוגית רב-מודאלית
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
מחקר זה בוחן התכנסות ייצוגית רב-מודאלית. ההיפותזה היא שרשתות נוירונים מתכנסות לייצוג משותף. המחקר מראה כי הראיות לכך חלשות יותר מקודמו. הייצוגים הרב-מודאליים חולקים מבנה סמנטי גס, אך לא נמצאה התכנסות עדינה.
תקציר מקורי באנגליתarXiv:2604.18572v3 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis posits that neural networks trained on different modalities (e.g., text and images) converge toward a shared representation of reality. If true, this has significant implications for whether modality choice matters at all. In this paper, we show that the evidence for this claim is substantially weaker than subsequent work suggests. The mutual $k$-nearest-neighbor metric used on 1024 text-image pairs in the original study captures only coarse structure. To keep the alignment from collapsing as one scales up the data, $k$ has to grow proportionally, undercutting the argument for fine-grained representational convergence. The reported increase in alignment with language model strength saturates fo
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית