יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

בנצ'מרק תרגום אחרון

Last Translation Benchmark
הבנצ'מרק האחרון לתרגום הוא אוסף של דוגמאות אנושיות וביקורתיות ששוברות את המודלים המובילים של תרגום מכונה. הוא כולל גם גישה חדשה להערכה, עם כללים מוגדרים לאימות כישלונות.
תקציר מקורי באנגליתarXiv:2609.04173v2 Announce Type: replace Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, opaque, and vulnerable to reward-hacking. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models.
קרא במקור המקורי