יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

הקשר מטעה: מדדים חדשים לאשליות תמיסות במודלים לשוניים

When Context Misleads: Surprisal, Energy and Attention Entropy as Metrics of Coherence Illusions in LLMs
חוקרים בדקו את התנהגותם של מודלים לשוניים הולנדיים ורב-לשוניים ביחס לאשליות תמיסות. הם מצאו שמודלים אלו מוטעים על ידי הקשר קודם, ופיתחו מדדים חדשים למדידת התופעה.
תקציר מקורי באנגליתarXiv:2606.21203v2 Announce Type: replace Abstract: Psycholinguistics studies show that human readers fall for coherence illusions: an incoherent discourse can seem coherent simply because a distractor matches what comes next. We investigate whether Dutch language models (6 monolingual and 4 multilingual) show the same behavior on texts that link back to earlier context with words such as 'again' and 'too'. First, we find that surprisal at the critical word tracks human acceptability judgments and eye-tracking data. Models are more surprised by incoherent continuations, but a matching distractor in the prior context reduces this surprisal. Second, attention entropy identifies heads that behave differently under coherence vs. incoherence. We find that ablating these heads shows transfer eff
קרא במקור המקורי