יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

פרטורבציה: עקבות אדוורסריאלי פשוט ויעיל ללמידת ייצוג במודלים שפה

Perturbation: A simple and efficient adversarial tracer for representation learning in language models
חוקרים הציגו שיטה חדשה ללמידת ייצוג במודלים שפה, המבוססת על פרטורבציה. השיטה מאפשרת לחשוף מבנים מועברים במודלים. היא פועלת על ידי ערבוב מודל שפה עם דוגמה אדוורסריאלית ומדידת ההשפעה על דוגמאות אחרות.
תקציר מקורי באנגליתarXiv:2603.23821v2 Announce Type: replace-cross Abstract: Linguistic representation learning in deep neural language models (LMs) has been studied for decades, but finding representations in LMs remains an unsolved problem. On the one hand, unconstrained alignments may trivialize the notion of representation (Sutter et al., 2025); on the other, even recently popularized linear approaches may not always be faithful to natural model behavior (Arora et al. 2024). Here we escape this dilemma by reconceptualizing representations not as patterns of activation but as conduits for learning. Our approach is simple: we perturb an LM by fine-tuning it on a single adversarial example and measure how this perturbation "infects" other examples. Perturbation makes no geometric assumptions, and unlike oth
קרא במקור המקורי