יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

זיהום מידע מנפח ציונים אך נדירות משנה דירוגים

Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards
מחקר חדש מראה כי זיהום מידע במודלים גדולים של שפה (LLM) מנפח ציונים אך נדירות משנה את הדירוגים. המחקר בדק 47 מודלים ציבוריים ו-74 מודלים שעברו עידון עם זיהום מידע ידוע, ומצא כי הזיהום משפיע מעט על הדירוגים. המחקר משתמש ב-LangChain ומודל LLaMA.
תקציר מקורי באנגליתarXiv:2609.02899v1 Announce Type: new Abstract: Benchmark contamination, the leakage of test items into training data, is widely described as a threat to the reliability of large language model (LLM) leaderboards. We argue that this concern conflates two distinct questions: whether contamination inflates absolute scores, and whether it reorders the ranking of models. We recast contamination as a violation of anchor-item invariance and measure it through the differential functioning of original versus semantically equivalent paraphrased items, a within-item contrast that holds the measured skill fixed and isolates memorization from capability. Using per-instance responses from 47 publicly released models and 74 models finetuned with a known dose of contamination, across four benchmarks (ARC
קרא במקור המקורי