כתבה
arXiv cs.CL ·
From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages
תקציר מקורי באנגליתarXiv:2607.24542v1 Announce Type: new Abstract: Automatic Text Recognition (ATR) now supplies digital humanities with large volumes of unstructured, heterogeneous, and often noisy text in ancient languages. Downstream semantic analysestext reuse identification, alignment, and semantic search-rely on sentence embeddings, yet existing methods transfer poorly to ancient languages: generic multilingual encoders underperform, specialized language models yield anisotropic representation spaces, and labeled similarity data is unavailable. We study two fully unsupervised strategies - TSDAE and contrastive sentence embedding (CSE) - that adapt a specialized token-level language model into a corpus-specific sentence encoder using only raw sentences. On the philologically central case of biblical reu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית