יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

איזה ניקוד? חידוש: ניקוד שאלות רב-בחירה עם תפריט תשובות

Your Benchmark Is Not Saturated: Reviving Multiple-Choice Evaluation with Answer Pooling
במאמר זה, המחברים מציגים חידוש בניקוד שאלות רב-בחירה. הם מציעים לשאול את המודל לזהות את התשובה הנכונה מתוך תפריט של תשובות של שאלות שונות. זאת כדי למנוע ניקוד קל ולהקשות על המודל. המחברים מציגים תוצאות של ניסויים שהראו שהחידוש יעיל.
תקציר מקורי באנגליתarXiv:2609.37494v1 Announce Type: new Abstract: Multiple-choice benchmarks are cheap to grade and are running out of room, and the standard remedy, writing harder items, is slow and repeated for every benchmark. A saturated benchmark still holds a harder task. Each question's wrong options are written for that question alone, so a model can score by eliminating a few options. We propose AnswerPool: take $N$ questions that share a context, pool all their options into one list, and ask the model to assign every question its answer. No item is written and no label changes. The chance of guessing a group right falls from $10^{-3}$ to $5\times10^{-7}$ for five four-option questions, and a model that recognizes its answers keeps its multiple-choice score, so the accuracy lost to pooling measures
קרא במקור המקורי