כתבה
arXiv cs.CL ·
Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora
תקציר מקורי באנגליתarXiv:2605.22660v2 Announce Type: replace Abstract: Moral language is subtle and culturally variable, making it difficult to translate faithfully across languages. Idiomatic expressions, slang, and cultural references introduce hard-to-avoid translation artifacts. Yet automated moral classification depends on language-specific annotated corpora that exist almost exclusively in English. We investigate whether LLM-based translation can bridge this gap, taking Polish as a test case. Using $\sim$50k morally annotated social media posts from a diverse range of topics, we apply a principled four-method validation pipeline: LaBSE cross-lingual embedding similarity, Centred Kernel Alignment (CKA), LLM-as-judge evaluation, and deep learning classifier parity tests. We show that despite shortcomings
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית