כתבה
arXiv cs.AI ·
SciDocBench: בנצ'מרק להבנת מסמכים מדעיים
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
SciDocBench הוא בנצ'מרק להבנת מסמכים מדעיים. הוא מורכב מ-124 שאלות שנוצרו על ידי מומחים ומחולקות ל-7 קבוצות יכולות ו-19 תת-משימות. SciDocBench מאפשר הערכה של מודלים רב-מודאליים.
תקציר מקורי באנגליתarXiv:2609.05141v1 Announce Type: new Abstract: Scientific papers require models to reason jointly over text, equations, figures, tables, code, and datasets while preserving the provenance of supporting evidence. Existing benchmarks typically evaluate these capabilities in isolation, leaving unclear whether multimodal models can support realistic scientific-reading workflows. We introduce SciDocBench, a workflow-centered benchmark for scientific document understanding. It contains 124 expert-authored and difficulty-screened questions organized into seven research-assistant capability groups and 19 subtasks across five scientific domains. Each question is instantiated under four matched conditions combining English or Chinese questions with all-images-first or interleaved document represent
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית