יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

למידה לקבלת תוצאות באמצעות רכיבת השפעה במרחב האמצעי

Learning to Retrieve via Reinforcement Learning in Embedding Space
אנו מציגים פרקטיקה של רכיבת השפעה במרחב האמצעי, המאפשרת למודלי קבלת תוצאות ללמוד לקבל תוצאות יעילות. ניתן להשתמש בפרקטיקה זו עם מודלי Gemini ו-GPT-5.
תקציר מקורי באנגליתarXiv:2610.07731v1 Announce Type: cross Abstract: Dense retrieval models are typically trained with contrastive objectives that learn effective representations but do not directly optimize retrieval metrics or downstream task performance. To address this problem, we introduce RELER (REinforcement LEarning for Retrieval), a reinforcement learning framework that enables existing embedding models to learn to retrieve directly in embedding space and align to task-specific rewards. We train RELER by sampling unit-length query and document embedding actions from von Mises-Fisher (vMF) distributions centered on normalized encoder outputs, scoring the resulting retrieval or downstream outcomes as rewards, and updating the encoder with REINFORCE using a leave-one-out baseline (RLOO). As exploration
קרא במקור המקורי