יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

דגמי תגמול משפטיים

Building Legal Reward Models for Grounding and Abstention
חוקרים פיתחו שיטה חדשה ליצירת דגמי תגמול משפטיים. המחקר מתמקד בשיפור היכולת של מודלים להסתמך על ראיות בתהליכי קבלת החלטות. השיטה החדשה מאפשרת למודלים להעריך טוב יותר את הראיות ולקבל החלטות מדויקות יותר.
תקציר מקורי באנגליתarXiv:2609.14739v1 Announce Type: cross Abstract: Large language models are increasingly used in high-stakes domains such as law, where systems must ground their reasoning in retrieved evidence and abstain when that evidence is insufficient. However, existing reward models are largely optimised for general preferences rather than contextual grounding, limiting their ability to evaluate these behaviours in retrieval-augmented generation (RAG) settings. We introduce a framework for transforming existing legal QA datasets into contextual preference data and use it to construct LegalRewardBench (LRB), a benchmark for evaluating grounded legal generation under noisy and insufficient retrieval conditions. Across general and legal contextual evaluation, we find that contextual DPO improves ground
קרא במקור המקורי