יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

תגמול בנתיב: למידת אותות השגחה ביניים לשאילתות תשובות גרפיות

Reward on Path: Learning Intermediate Supervision Signals for Knowledge Graph Question Answering
RoP הוא כלי ללמידת אותות השגחה ביניים לשאילתות תשובות גרפיות. הוא משתמש באותות השגחה מותאמים לשאילתות ומשפר את התוצאות. המחקר מראה שיפור של 2.6% ב-F1.
תקציר מקורי באנגליתarXiv:2605.10791v2 Announce Type: replace Abstract: Knowledge Graph Question Answering (KGQA) aims to answer user questions by reasoning over Knowledge Graphs (KGs). Recent methods use supervision derived from answer labels or refined by Large Language Models (LLMs) to train models that retrieve KG evidence for LLM-based answer reasoning. However, answer-derived supervision treats every answer-reaching path as correct and thus yields noisy training signals, whereas LLM-refined supervision mitigates this noise at substantial cost. To address these limitations, we propose Reward on Path (RoP), a framework to learn a lightweight, question-conditioned path reward from answer labels with an asymmetric objective. Paths reaching the same answer are supervised jointly as a bag, allowing the reward
קרא במקור המקורי