יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

עדות צידדה, אך לא נבחנה: תיקון נגד-סיבתי של נאמנות רשימת-הגירעין המשפטי

Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness
מודלי שפה גדולים נוטים להציג הסברים משפטיים שאינם אמינים. המאמר עורך תיקון נגד-סיבתי על נאמנות רשימת-הגירעין המשפטי של מודלי שפה גדולים. המחקר מצא כי כאשר נדרשים המודלים להצדיק פסק דין על ידי ציון הסמכות המשפטית שעליה הוא מבוסס, הם צועדים ב-66.7%-100% של הפקות, בעוד שהפסק הדין נשאר חסר תקדים כאשר הסמכות המשפטית נשתנה. המחקר גם מצא כי המודלים נוטים להצדיק פסק דין על ידי ציון סמכות משפטית שאינה קשורה לו, וכי הפסק הדין נשאר חסר תקדים כאשר הסמכות המשפטית נשתנה.
תקציר מקורי באנגליתarXiv:2610.12361v1 Announce Type: cross Abstract: Large language models increasingly justify legal decisions by naming the statute or precedent behind a verdict, treated as evidence that the decision follows from it. We test this directly: holding case facts fixed, we substitute the named legal authority for an unrelated one and decode a model's evolving verdict from its hidden states. Across seven open-weight models (8B-70B) and four benchmarks spanning judicial and contractual reasoning, when explicitly required to justify a verdict by naming the governing authority, models name the correct one in 66.7%-100% of generations, while the verdict changing when the authority changes is far less consistent: 0.0%-21.7% on CaseHOLD, 30.0%-76.7% on ECHR and SCOTUS, and 43.3%-50.0% on ContractNLI.
קרא במקור המקורי