יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

MELLA: גישור בין יכולת לשונית לרקע תרבותי

MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs
MELLA הוא מאגר נתונים רב-מודאלי לשפות בעלות משאבים מוגבלים. הוא משלב זוגות תמונה-טקסט ילידים עם תיאורי תמונות מתורגמים כדי לשפר את היכולת הלשונית והרקע התרבותי. MELLA זמין ב-https://opendatalab.com/applyMultilingualCorpus.
תקציר מקורי באנגליתarXiv:2508.05502v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) perform strongly in high-resource languages, yet often produce fluent but culturally "thin" descriptions in low-resource settings. We argue that this failure is not merely a linguistic limitation: culture-specific visual knowledge depends on native visual-textual alignments that translation-centric pipelines rarely provide. We present MELLA, a multimodal dataset across eight low-resource languages, designed to support linguistic fluency and cultural groundedness. MELLA uses a dual-source strategy that combines native web image-alt-text pairs for culture-grounded supervision with generated-and-translated image descriptions for linguistically rich supervision, explicitly separating two learning
קרא במקור המקורי