כתבה
arXiv cs.CL ·
הפיכת טקסט קלאסי לוויקטור: למידת ייצוג לזוגות שיר-פירוש
Vectorizing Classical Tamil: Representation Learning for Verse-Commentary Pairs
במאמר זה, נבנה קורפוס של 1,262 זוגות שיר-פירוש מן הטמילית הקלאסית, ונבחן את יכולת הלמידה של רכיבי הייצוג לחידוש זה.
תקציר מקורי באנגליתarXiv:2609.04755v1 Announce Type: new Abstract: We construct a corpus of 1,262 verse-commentary (urai) pairs from five Classical Tamil source sections, ranging from technical grammatical prose to modern paraphrase, and ask what information representation learning can recover. We train recurrent and Transformer encoders, a Siamese-style pair-matching network, an mBART-style encoder-decoder, and a decoder-only language model. Each analysis is interpreted against an appropriate control on the same data. TF-IDF provides a strong no-training lexical retrieval baseline, alongside representation analyses and generation controls for the learned models. A fixed string containing the 25 most frequent commentary words scores higher on generation overlap than the decoder-only model. Canonical correlat
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית