יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

LexAgentHallu: בניית תקן לפרופיל תקלות הזיה באגנטים משפטיים

LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents
נחשפה תקלה באגנטים משפטיים: LexAgentHallu, תקן חדש לפרופיל תקלות הזיה. התקן כולל 3414 תצוגות ו-17 קטגוריות משפטיות. התקן נבנה דרך פייפלין של ארבעה שלבים, וכולל 27 תת-קטגוריות של תקלות הזיה. התקן נבדק על 18 אגנטים פרטיים ופתוחים.
תקציר מקורי באנגליתarXiv:2609.09754v1 Announce Type: new Abstract: As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic hallucinations where tool-call and reasoning errors cascade into fabricated holdings and miscited authority. However, existing legal benchmarks evaluate only single-turn QA with outcome-level metrics, while agentic hallucination benchmarks lack legal-specific diagnostic capability. Neither answers to what extent and how a legal agent hallucinates along its trajectory. To address these limitations, we introduce LexAgentHallu, a legal agentic hallucination benchmark designed to evaluate to what extent and how legal agents fail along multi-step trajectories. Built through a four-stage expert-in-the-loop pipeline, LexAgentHallu contains 3414
קרא במקור המקורי