כתבה
arXiv cs.AI ·
ArgGYM: בנך לבדיקה פרוצדורלי לתמיכה בהיגיון סביל
ArgGYM: A Procedural, Engine-Verified Benchmark for Structured Defeasible Reasoning
ArgGYM הוא בנך לבדיקה פרוצדורלי לתמיכה בהיגיון סביל. הוא מכיל 1,440 מקרי בדיקה מאומתים ומאפשר אימון מודלים באמצעות תגמולים אוטומטיים. הבנך מורכב מ-12 משימות ומשתמש במנוע טיעונים סמלי לחישוב מצבים פורמליים.
תקציר מקורי באנגליתarXiv:2609.38409v1 Announce Type: new Abstract: Recent progress in large language model reasoning has been driven by benchmarks and reinforcement learning environments with automatically verifiable rewards, particularly in mathematics, code, and formal logic. These settings make model accuracy easier to evaluate and optimize, but it remains unclear how far success under fixed problem specifications and stable evaluation criteria transfers to reasoning outside such domains. Real-world reasoning often proceeds under incomplete and revisable information: conclusions may be supported provisionally, defeated by counter-evidence, reinstated by further arguments, or revised when stronger reasons become available. Reasoning of this kind is generally referred to as defeasible reasoning. We introduc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית