יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

FrontierChallenge: בדיקת השלמת עבודות מדעיות

FrontierChallenge: Evaluating Scientific Workflow Completion
FrontierChallenge הוא בנץ'מרק הבודק השלמת עבודות מדעיות. הוא כולל 300 עבודות מדעיות שונות, ונבדקו 97 מהן. התוצאות הראו שאף מודל לא הצליח להשלים יותר מ-20 עבודות. Claude Code היה אחד המודלים שנבדקו.
תקציר מקורי באנגליתarXiv:2608.24979v2 Announce Type: replace Abstract: Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures
קרא במקור המקורי