כתבה
arXiv cs.AI ·
Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora
תקציר מקורי באנגליתarXiv:2601.14994v2 Announce Type: replace-cross Abstract: Data contamination can invalidate benchmark evaluation when a model benefits from memorized evaluation content rather than genuine generalization. Yet contamination is difficult to audit when the exposed content differs in language from the evaluation benchmark. We study this failure mode by deliberately exposing four open-weight instruction-tuned LLMs to Arabic translations of MMLU and XQuAD evaluation items at increasing exposure levels, then evaluating them on the original English tasks. This controlled setup is a proxy for contamination rather than a reconstruction of real-world pretraining leakage. We first test two English-centric post-hoc probes, TS-Guessing and Min-K%++, and find that their signals largely disappear under tr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית