יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

LLMs: האם יכולים לעקוב אחר הלוגיקה הרפואית המקצועית? במבחן לתיקון-סדר לוגי ברמה הגבוהה

Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment
במבחן חדש, חוקרים בדקו את יכולת ה-LLMs לעקוב אחר הלוגיקה הרפואית המקצועית. התוצאות? חלק מה-LLMs נכשלו בביצוע תיקון-סדר לוגי ברמה הגבוהה.
תקציר מקורי באנגליתarXiv:2609.11185v1 Announce Type: new Abstract: Evidence-based medicine demands strict logical consistency, yet current evaluations of large language models (LLMs) prioritize superficial label matching over genuine reasoning. We introduce LogiMed-RoB, a benchmark grounded in Cochrane Risk of Bias (RoB) 2.0 expert logic, comprising 860 randomized controlled trials (RCTs) and 14,820 queries. It evaluates models under the Hierarchical Logical Consistency (HLC) framework across four dimensions: Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness. Experiments on 10 state-of-the-art LLMs reveal a catastrophic Error Compounding Effect: despite the top model reaching 98.88% Atomic Consistency, its end-to-end consistency collapses to 45.13%, with several ope
קרא במקור המקורי