יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

הערכה חסרת התייחסות של תהליכי היגיון בשאלות פתוחות

Reference-Free Evaluation of Reasoning in Open-Ended Question Answering
חוקרים הציגו שיטה חדשה לבדיקת תהליכי היגיון של מודלים כמו LLM. השיטה משתמשת ב-NLI כדי לוודא את תוצאות המודל. היא נבדקה במסגרת שאלות מתמטיות ורפואיות.
תקציר מקורי באנגליתarXiv:2607.19678v1 Announce Type: new Abstract: AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rather than a single final answer. We propose a reasoning-based, reference-free framework for auditing LLM-generated outputs. The method decomposes a generated reasoning trace into segments, labels local premise-target relations using Natural Language Inference (NLI), and organizes these relations into a hypergraph. A deterministic backward AND-OR search then assigns segment-level audit labels that indicate how each segment is grounded within the generated response. We evaluate the framework in two settings: deductive mathematical reasoning with Hard2Verify, and open-ended medical reasoning with UroReason, a
קרא במקור המקורי