יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

דירוג ותיקון: טכניקה אדפטיבית לבדיקת LLM עם ציונים רציפים

Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
אוריינטציה אדפטיבית לבדיקת LLM עם ציונים רציפים. טכניקה זו מאפשרת דירוג ותיקון יעיל של LLM על ידי שימוש בציונים רציפים. הטכניקה נבחנה על חמש בסיסי נתונים שונים והראתה תוצאות משמעותיות.
תקציר מקורי באנגליתarXiv:2601.13885v2 Announce Type: replace-cross Abstract: Computerized Adaptive Testing (CAT) has proven effective for efficient LLM evaluation on multiple-choice benchmarks, but modern LLM evaluation increasingly relies on generation tasks where outputs are scored continuously rather than marked correct/incorrect. We present a principled extension of IRT-based adaptive testing to continuous bounded scores (ROUGE, BLEU, LLM-as-a-Judge) by replacing the Bernoulli response distribution with a heteroskedastic normal distribution. Building on this, we introduce an uncertainty aware ranker with adaptive stopping criteria that achieves reliable model ranking while testing as few items and as cheaply as possible. We validate our method on five benchmarks spanning n-gram-based, embedding-based, an
קרא במקור המקורי