כתבה
arXiv cs.CL ·
SEA-CLIP-Tiny: מודל עיבוי טקסט-חזון רב-לשוני יעיל
SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages
SEA-CLIP-Tiny הוא מודל עיבוי טקסט-חזון רב-לשוני קומפקטי לדרום-מזרח אסיה. המודל משיג ביצועים חזקים באיחזור תמונות-טקסט. הוא משתמש במסגרת CLIP-KD והדרכת מורה רב-לשונית.
תקציר מקורי באנגליתarXiv:2609.30739v2 Announce Type: replace Abstract: Multilingual text-vision embedding models are essential for cross-lingual image-text retrieval, but Southeast Asian languages remain poorly supported due to the region's linguistic diversity and limited data and computing resources. In this paper, we introduce SEA-CLIP-Tiny, a compact multilingual text-vision embedding model for Southeast Asia with fewer than 50M parameters. Our model adapts a CLIP-KD-style framework to Southeast Asian multilingual settings through regional data curation and multilingual teacher guidance. Experiments across seven Southeast Asian languages show that SEA-CLIP-Tiny achieves the strongest average retrieval performance among the evaluated student models, reaching 12.9%, 31.5%, and 42.2% at R@1, R@5, and R@10,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית