יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Doc2FRC: תרגום מסמכים באמצעות חלוקה לחלקים קבועים

Doc2FRC: Length-Consistent Document-Level Machine Translation via Fixed-Range Chunking
Doc2FRC משפר תרגום מסמכים על ידי חלוקה לחלקים קבועים, מודל LLM גדול עם חלונות הקשב הארוכים משפר את איכות התרגום. השיטה מומלצת לתרגום מסמכים ארוכים.
תקציר מקורי באנגליתarXiv:2609.12674v1 Announce Type: new Abstract: Advanced large language models (LLMs) with long context windows can substantially reduce input truncation in document-level machine translation (DocMT). However, direct Doc2Doc translation remains prone to n-gram repetition and progressive quality degradation. A common remedy is to segment the document into finer-grained chunks. Nonetheless, conventional rule-based chunking approaches fail to handle the length distribution mismatch between training and inference. To address this, we introduce Fixed-Range Chunking (FRC), utilizing dynamic programming to partition documents into chunks within a predefined length interval. By consistently applying FRC during training and inference, the input documents of any length are mapped to the same length
קרא במקור המקורי