כתבה
arXiv cs.CL ·
אימות ותחושת הסתברות: מחקר גדול-ממד של ביצועי MCQA בעברית וב-22 שפות נוספות
Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs
במחקר זה, נבחנו שיטות לאימות ותחושת הסתברות ב-LLMs ב-22 שפות, כולל עברית. נמצא כי פקודות הגברה באנגלית תורמות לשיפור בביצועי האימות, וכי סקאלת המודל תורמת לבחירת שיטת האימות הטובה ביותר.
תקציר מקורי באנגליתarXiv:2607.06327v3 Announce Type: replace Abstract: Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English. We present the first large-scale evaluation of UE methods across 22 languages, spanning high-, mid-, and low-resource settings. Using two human-curated Q&A datasets, we compare open and closed box UE methods (nine in total) across different model sizes and architectures while eliciting long-form reasoning, avoiding LLM-as-a-judge and embedding-based scoring, which can introduce evaluation noise. We report three main actionable findings. First, we find that prompting models to reason in English while keeping questions in low-resource languages substantially improves UE performance, suggesting tha
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית