כתבה
arXiv cs.CL ·
Cross-Lingual Alignment for Decoder-Only Models using MoE Routers
תקציר מקורי באנגליתarXiv:2610.01921v2 Announce Type: replace Abstract: Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, higher cross-lingual representational alignment leads to improved cross-lingual transfer. In this paper, we propose a novel approach to reimagine cross-lingual contrastive learning given the architectural constraints of modern LLMs. Rather than applying an auxiliary alignment loss on hidden states, we propose using the outputs of the mixture-of-experts (MoE) routers as the target for alignment. Router outputs lend themselves better to pooling over
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית