יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

בקשה לידיד עתיק: הרכבה והתמודדות עם תסכול זמני בשאלות תשובה סטטוטוריות באמצעות LLM

Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering
במאמר זה, חוקרים חוקרים שני תסכולי זמן בשאלות תשובה סטטוטוריות באמצעות LLM: תסכול קפיצה-אחר-קפיצה והטיה זמנית. הם מציגים בנק אקספרט-מאומת, ובודקים את חמשה LLM של OpenAI, Anthropic ו-DeepSeek.
תקציר מקורי באנגליתarXiv:2605.23497v2 Announce Type: replace Abstract: Large language models are increasingly used for legal research, yet their fixed training cutoffs and reliance on static parametric knowledge are at odds with the evolving nature of statutory law. We study two temporal failure modes: post-cutoff staleness, where models apply superseded rules after legislative amendments, and recency bias, where models prefer newer provisions even when a historical version governs the fact pattern. To this end, we present a benchmark of 312 expert-validated, time-sensitive German statutory QA pairs spanning three categories: Post-Cutoff Amendment Questions, Pre-Amendment Questions, and Multi-Provision Pre-Amendment Questions. We evaluate five LLMs by OpenAI, Anthropic and DeepSeek under four inference setti
קרא במקור המקורי