יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Sci-MMR: בנצ'מרק לתיאום ראיות מדעיות

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
Sci-MMR הוא בנצ'מרק לבדיקת סוכנים רב-מודאליים. הוא בודק יכולתם לחפש ראיות, לנתח נתונים וליצור השערות מדעיות. הבנצ'מרק כולל 235 משימות בארבע תחומים מדעיים.
תקציר מקורי באנגליתarXiv:2609.11243v1 Announce Type: new Abstract: Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-step evidence grounded reasoning that progressively acquires, integrates, and verifies evidence before reaching a conclusion. Existing multimodal benchmarks, however, largely evaluate final-answer accuracy, leaving open whether predictions are actually supported by traceable scientific evidence. We introduce Sci-MMR, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions. Sci-MMR comprises 235 multi-hop reasoning tasks spannin
קרא במקור המקורי