יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

משיכים יעילים לשפות פולניות ואירופיות

Parameter-Efficient Retrievers for Polish and European Languages
פותחו משיכים קומפקטים לשפות פולניות ואירופיות. PolDense ו-EuroDense הם שני מודלים שפותחו, התומכים בהקשרים של עד 8,192 תווים. המודלים הוכשרו באמצעות צינור בן שלושה שלבים.
תקציר מקורי באנגליתarXiv:2609.12913v1 Announce Type: new Abstract: Dense retrieval systems increasingly rely on multi-billion-parameter language models, whose memory and computational requirements make large-scale indexing, frequent corpus updates, and low-latency serving costly. We present a three-stage training pipeline for developing compact and efficient retrievers that remain competitive with substantially larger models. The pipeline combines cross-lingual alignment, relational knowledge distillation, and contrastive fine-tuning. It requires no original ground-truth relevance labels, relying exclusively on supervision generated by strong embedding models and rerankers utilised as teachers. Using this pipeline, we develop PolDense and EuroDense, both supporting contexts of up to 8,192 tokens. PolDense is
קרא במקור המקורי