כתבה
arXiv cs.CL ·
Improving Cross-Lingual Token Representations by Adding a Pinch of SALT
תקציר מקורי באנגליתarXiv:2609.09953v1 Announce Type: new Abstract: Cross-lingual sentence encoders enable scalable transfer across hundreds of languages, powering applications such as translation mining and zero-shot learning in low-resource settings. Although trained for sentence-level alignment, they are increasingly also applied to token-level tasks such as hallucination detection and sequence tagging, exposing a mismatch between training and usage. We propose SALT, a lightweight post-training method that improves token representations by injecting span-level supervision into existing sentence encoders. Across five multilingual token-level benchmarks, SALT achieves the best overall results on four of them, outperforming alternative fine-tuning strategies and competitive encoders. It also improves sentence
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית