יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אמינות בסוכנויות מחקר משפטי

Legal Research Bench: Measuring End-to-End Reliability in Long-Horizon Legal Research Agents
חוקרים פיתחו את Legal Research Bench, בנך' לבדיקת אמינות בסוכנויות מחקר משפטי. הבנך' בודק 13 מודלים, כולל Claude, ומוצא שאף אחד מהם אינו אמין לחלוטין.
תקציר מקורי באנגליתarXiv:2610.00609v1 Announce Type: new Abstract: Legal research is a core and time-consuming legal workflow. Lawyers must identify controlling authority, verify that it remains valid, reconcile statutes and cases, and synthesize a grounded answer. Language model agents are a natural fit for this retrieval-intensive workflow, and automating even part of it would be valuable. But that value depends on reliability: a single missing authority, stale citation, or wrong legal conclusion can make an otherwise plausible answer unusable. We introduce \textbf{Legal Research Bench} (LRB), a benchmark of 413 open-ended U.S. legal research questions written by experts, each paired with a gold answer, supporting authorities, and a binary grading rubric. We evaluate thirteen frontier models in a harness w
קרא במקור המקורי