יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

SalamandraTA ב-WMT 2026: דוגמאות קשות הן מורים טובים יותר

SalamandraTA at WMT 2026 Terminology Shared Task: Hard Examples Are Better Teachers
SalamandraTA השיגה תוצאות טובות יותר במשימת WMT26 על ידי ניקוי נתונים. המערכת, SalamandraTA-7b-instruct v3.0, כוללת נתונים שנבנו בצירוף סינתטי ומופעלת על ידי פייפלינג תיאורי. הפייפלינג כולל תרגום תאורטי, תרגום תאורטי-סינתטי, ופייפלינג תיאורי-סינתטי.
תקציר מקורי באנגליתarXiv:2609.09999v1 Announce Type: new Abstract: Terminology-aware translation asks for more than a correct translation: the output must use the exact terms a glossary prescribes. The standard recipe, fine-tuning on glossary-annotated translation pairs, hides an inefficiency: for most examples the glossary prescribes exactly what the model would have produced anyway, so they teach nothing about following a glossary. We therefore keep only the examples where the model's own translation contradicts the glossary. In a controlled study at fixed data volume, this selection alone raises term accuracy from 78.7% to 89.9%. The filtered data, built by a two-way synthetic pipeline on open models, is part of the instruction-tuning mixture of our public release SalamandraTA-7b-instruct v3.0, which, use
קרא במקור המקורי