כתבה
arXiv cs.CL ·
איתור נוירונים של שפה בערך-שפה גדול: הכרה תפוצה-מודעה
Distribution-aware Language Neuron Identification in Multilingual Large Language Models
איתור נוירונים של שפה בערך-שפה גדול: חידוש חדש נוסה לזהות נוירונים ספציפיים לשפה בערכים-שפה גדולים. החידוש נועד לשפר את יכולת המערכת להבין שפות שונות.
תקציר מקורי באנגליתarXiv:2609.10993v1 Announce Type: new Abstract: Multilingual large language models (mLLMs) contain a small fraction of feed-forward neurons that are sensitive to particular languages, commonly termed language-specific neurons. Existing work measures language specificity using the entropy of each neuron's language-wise probabilities of being active, where a neuron is considered active when its activation value is positive. However, this approach may not fully capture the multilingual nature of mLLMs, where language representations are distributional and mutually related. We propose Distribution-aware Language Neuron selection, which leverages pairwise relationships between per-language activation distributions over the full activation range, including negative values. Specifically, we quant
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית