כתבה
arXiv cs.CL ·
בדיקה על פי הכריכה: טיהור סטטיסטיקות LLM למניעת פריצת פרטים פנימיים
Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage
טיהור סטטיסטיקות LLM למניעת פריצת פרטים פנימיים. נמצאו פריצות פרטים בבסיס נתונים TruthfulQA.
תקציר מקורי באנגליתarXiv:2609.13003v1 Announce Type: new Abstract: Binary-choice truth benchmarks ask models to choose between a correct and an incorrect answer, but if the two answers differ systematically in surface-level features, models can exceed chance without performing the intended reasoning. We show that this failure mode is detectable and can be exploited by downstream classifiers. In TruthfulQA, a simple six-feature logistic classifier achieves substantial accuracy in separating correct from incorrect answers. We further show that similar surface-level artifacts are present in additional benchmarks. To counteract this, we developed a general mechanism to clean them by removing the most leakage-reinforcing pairs. We release a version of TruthfulQA with surface-feature leakage reduced close to chanc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית