כתבה
arXiv cs.AI ·
כאשר הציטציות מטעים? בנקאות תקן לזיהוי הלקוחות הקשישים של הטקסט
When Citations Mislead? A Claim-Level Benchmark for Legal Hallucination Detection
במאמר זה, נוצרה בנקאות תקן לזיהוי הלקוחות הקשישים של הטקסט. הבנקאות נבנתה על ידי 3,396 טענות פרנטז'י-סגנון, שנותרו כטענות תקניות, נדחו או לא נמצאו. המאמר נותן תוצאות טובות למודלי LLM בשביל הזיהוי של הטענות הקשישים.
תקציר מקורי באנגליתarXiv:2610.10971v1 Announce Type: cross Abstract: Large language models are increasingly used in legal research and drafting, but they can still produce claims that sound convincing without being supported by the cited source. We introduce PARCEL, a benchmark for checking whether a legal claim is supported by the underlying authority. Using recent New York State Court of Appeals decisions, we build a dataset of 3,396 parenthetical-style claims labeled as Supported, Refuted, or Not Found. We cast this task as a three-way natural language inference problem and evaluate several state-of-the-art LLMs in a zero-shot setting. Although the strongest models reach up to 0.97 accuracy, the results also show an important weakness: models still incorrectly mark unsupported claims as supported, even wh
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית