כתבה
arXiv cs.CL ·
Building Legal Reward Models for Grounding and Abstention
תקציר מקורי באנגליתarXiv:2609.14739v1 Announce Type: new Abstract: Large language models are increasingly used in high-stakes domains such as law, where systems must ground their reasoning in retrieved evidence and abstain when that evidence is insufficient. However, existing reward models are largely optimised for general preferences rather than contextual grounding, limiting their ability to evaluate these behaviours in retrieval-augmented generation (RAG) settings. We introduce a framework for transforming existing legal QA datasets into contextual preference data and use it to construct LegalRewardBench (LRB), a benchmark for evaluating grounded legal generation under noisy and insufficient retrieval conditions. Across general and legal contextual evaluation, we find that contextual DPO improves grounded
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית