יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

תאוריה של ייצוגים פלאטוניים במודלי שפה

A theory of platonic representations in language models
אנו מציגים תאוריה של ייצוגים פלאטוניים במודלי שפה, המסבירה את הדמיון הבין-לשוני בשכבות הפנימיות של מודלי שפה רב-לשוניים.
תקציר מקורי באנגליתarXiv:2610.07168v1 Announce Type: cross Abstract: Representations of translated sentences are similar in the inner layers of multilingual language models -- an observation connected to the platonic representation hypothesis, yet unexplained theoretically. We provide an explanation based on the assumption that data have a hidden hierarchical structure whose abstract levels are shared across languages while surface levels are modality- or language-specific. Concretely, we generate synthetic languages from probabilistic context-free grammars sharing upper-level but not lower-level production rules. In this setting the Bayes-optimal next-token predictor is belief propagation (BP); encoding its messages in successive layers yields analytical predictions that agree well with transformers trained
קרא במקור המקורי