יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

הערכת מחקר אוטומטי

How Is Automated Research Evaluated? A Survey of Benchmarks and Evaluation Practices
מחקר אוטומטי תומך בסינתזה של ספרות, רעיונות, ניסויים, כתיבה וביקורת עמיתים. המחקר בוחן את הדרכים להערכת מערכות אלו.
תקציר מקורי באנגליתarXiv:2610.11877v1 Announce Type: new Abstract: Automated research systems support literature synthesis, ideation, experiments, writing, and peer review, but their evaluation is dispersed across tasks, benchmarks, and studies that are difficult to compare directly. We review this literature from the perspective of evaluation design and evidence, covering six targets: literature synthesis, research ideation, executable workflows, scholarly writing and communication, automatic peer review, and end-to-end research. We compare task construction, evidence sources, evaluators, and scoring procedures to explain the capabilities assessed by different designs. Our synthesis highlights three recurring lessons: output checks, process checks, and human studies provide complementary information; evalua
קרא במקור המקורי