כתבה
arXiv cs.AI ·
Online Language Adaptive Sampling for Better Distributed Cross-lingual Gains
תקציר מקורי באנגליתarXiv:2609.14969v1 Announce Type: cross Abstract: Realignment is a promising approach for improving the cross-lingual transfer ability of multilingual language models, particularly for extremely low-resource languages (LRLs). However, existing realignment methods rely on uniform and random sampling of parallel sentences across languages, which may be suboptimal under limited batch sizes. In practice, models may benefit from seeing certain languages more frequently, especially those that are poorly aligned, and the optimal distribution can evolve throughout training. In this work, we propose a simple yet effective adaptive sampling strategy that assigns trainable sampling probabilities to each language. Languages that contribute more to the realignment loss are sampled more frequently in su
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית