כתבה
arXiv cs.CL ·
בסיס התרגום האחרון
Last Translation Benchmark
בסיס התרגום האחרון: בסיס תרגום חדש שבודק את גבולות המודלים המתקדמים. הבסיס כולל דוגמאות שבורות למודלי תרגום ובאים עם כלי ניפוי ידני.
תקציר מקורי באנגליתarXiv:2609.04173v1 Announce Type: new Abstract: For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, vulnerable to reward-hacking, and provide unactionable assessments. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking objective progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break lead
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית