כתבה
arXiv cs.CL ·
DeepResearch Bench II: ניתוח סוכני מחקר עמוקים באמצעות קריטריות מדויקות מדווחות על ידי דיווחי מומחים
DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports
במאמר זה נוצרה תקן חדש לבדיקת סוכני מחקר עמוקים באמצעות קריטריות שנכתבו על ידי מומחים. התקן כולל 132 תפקידי מחקר שונים ב-22 תחומים, וכל תפקיד דורש מהסוכן ליצור דיווחי מחקר שיוצקל על ידי 9,430 קריטריות דויקות. הקריטריות נבנו באמצעות פייפלין של LLM+אדם, והן נבנו על ידי כ-400 שעות של עריכה של מומחים. המחברים ניתחו כמה סוכני מחקר עמוקים מובילים על תקן זה ומצאו שאף על פי שהסוכנים החזקים ביותר עומדים בפחות מ-50% מהקריטריות, יש פער גדול בין הסוכנים העמוקים לבין המומחים. המחברים פרסמו את התקן, את הסקריפטים לבדיקה ואת כל הקריטריות באתר https://github.com/imlrz/DeepResearch-Bench-II.
תקציר מקורי באנגליתarXiv:2601.08536v3 Announce Type: replace Abstract: Deep Research Agents (DRA) aim to help users search the web, synthesize information, and deliver comprehensive investigative reports. Prior benchmarks often either under-evaluate a system's ability to produce meaningful insights and high-quality writing, or adopt coarse or LLM-defined criteria that are hard to verify and can diverge from human expert judgment. To address these issues, we introduce Deep Research Bench II, a new benchmark for evaluating DRAs. It contains 132 grounded research tasks across 22 domains; for each task, an agent must produce a research report that is evaluated by a set of 9,430 fine-grained binary rubrics in total, covering three dimensions: information recall, analysis, and presentation. All rubrics are derived
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית