כתבה
arXiv cs.CL ·
בערמה: בניית התייחסויות פסאודו לבדיקת MT
In the Blind: Building Pseudo-References for MT Evaluation
בניית התייחסויות פסאודו לבדיקת MT בלי התייחסויות אנושיות. נבנו התייחסויות פסאודו ל-7 מודלים, והתייחסויות אלה נבחרו על ידי 3 מודלי עריכה חינמיים. התייחסויות פסאודו נבחרו על ידי GPT-5.5. התייחסויות פסאודו נבנו על ידי 3,277 דוקומנטים רשמיים.
תקציר מקורי באנגליתarXiv:2609.13611v1 Announce Type: new Abstract: The WMT26 General MT task evaluates systems on 10 language pairs that have no human references (neither translated from scratch nor post-edited from MT output by humans). We describe how we built the pseudo-references for these pairs and six other language pairs (in which some forms of human references are available): seven models translate the 3,277 official documents under up to five prompt conditions, giving a total of 26 system-prompt combinations; then three reference-free quality estimation (QE) models score every candidate; and a per-document selector picks one translation, which GPT-5.5 post-edits where needed. Working without references exposed a failure mode of QE-guided selection: the metrics rank fluent output in the wrong languag
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית