כתבה
arXiv cs.AI ·
מדריך טכני לכלי למדידת התאמה-פרטית במודלי שפה מטען-מעבד
Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models
כלי חדש למדידת התאמה-פרטית במודלי שפה מטען-מעבד, כולל שמות מודלי שפה (GPT-5)
תקציר מקורי באנגליתarXiv:2609.05333v1 Announce Type: new Abstract: A transformer language model assigns a single, context-independent vector to a word type at its embedding layer, yet is widely believed to individuate that word's occurrences by context in its later layers. Testing this belief cleanly requires a construct that holds the word form fixed while its context and intended sense vary in a controlled, labeled way. This manual documents an open toolkit built around such a construct, which we call a bridge form: a single written word that recurs, unchanged, across two or more subject domains with a different sense in each. We describe, and justify, every stage of the pipeline: the declarative specification of bridge forms and their source domains, corpus acquisition from Wikipedia, occurrence localizat
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית