כתבה
arXiv cs.CL ·
MCQA עם אפשרות להימנעות
Penalty-Framed No-Valid-Option MCQA: Analyzing LLM Abstention under Invalid Choices
חוקרים את היכולת של מודלים גדולים לשפה (LLM) להימנע מבחירה כאשר אף אחת מהאפשרויות אינה נכונה. המחקר משתמש במתודולוגיה חדשה הנקראת penalty-framed no-valid-option MCQA, ומראה כי אפילו מודלים עם דיוק גבוה בשאלות מרובות בחירה (MCQA) עדיין עלולים לבחור אפשרות שגויה.
תקציר מקורי באנגליתarXiv:2610.08153v1 Announce Type: new Abstract: Multiple-choice question answering (MCQA) is commonly used to evaluate large language models under the assumption that one of the provided options is correct, typically using answer-selection accuracy. However, in real deployments, users or retrieval systems may provide invalid option sets in which none of the listed choices is correct, and selecting one of them may incur downstream cost. We study this setting as penalty-framed no-valid-option MCQA. Using the mathematics subset of MMLU-Pro, we remove the labeled correct option, allow models to either choose a remaining option or output ABSTAIN, and penalize invalid forced-choice responses. We further introduce correct-conditioned analysis, evaluating abstention only on instances that the mode
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית