כתבה
arXiv cs.CL ·
$M^2PO$: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation
תקציר מקורי באנגליתarXiv:2510.13434v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with human preferences is pivotal for Machine Translation (MT), yet current approaches are often hindered by misleading reward signals. Our analysis reveals that prevailing Quality Estimation (QE) models exhibit a systematic blind spot toward partial errors, specifically partial hallucinations and omissions, often favoring superficially fluent but unfaithful translations. To address this issue, we propose $M^2PO$ (Multi-Perspective Multi-Pair Preference Optimization), a data-centric framework for preference optimization in machine translation. First, to correct the bias toward fluency, $M^2PO$ uses a dual-perspective mechanism that decouples semantic fidelity from fluency and prioritizes faithfulness
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית