כתבה
arXiv cs.CL ·
TransClean: בנך לזיהוי תרגומים נקיים
TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model Outputs
TransClean הוא בנך לזיהוי תרגומים נקיים מרעשים. הוא בוחן 790,000 תרגומים מ-12 מודלים שונים, ומציע שתי שיטות להסרת רעשים. הבנך כולל 9,900 זוגות של תרגומים רועשים ונקיים.
תקציר מקורי באנגליתarXiv:2609.11399v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for machine translation, yet their outputs often contain additional text beyond the translation itself, such as language labels, explanations or bilingual repetitions, which we term translation noise. Despite its prevalence, this problem lacks dedicated benchmarks and systematic study. We analyze over 790,000 translation outputs from 12 LLMs across 22 language pairs (LPs) and identify 12 recurring noise patterns, which we group into formatting and content noise. Building on the observed patterns, we construct TransClean, a controlled benchmark of 9,900 pairs of noisy and clean translation outputs, comprising 8,800 synthetically generated instances and 1,100 manually curated authentic instance
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית