יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Mr.LHDR: בנצ'מרק לסוכנים עמוקים

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents
Mr.LHDR הוא בנצ'מרק לבדיקת סוכנים עמוקים במחקר ארוך טווח. הוא בודק יכולתם לבצע מחקר מורכב עם ראיות רב-מודאליות, כולל תמונות, מפות וסרטונים. התוצאות מראות כי אפילו המערכות החזקות ביותר משיגות רק 43.1% דיוק.
תקציר מקורי באנגליתarXiv:2609.11318v1 Announce Type: new Abstract: Deep research agents are increasingly capable of web search, tool use, multimodal evidence analysis, and information synthesis. However, existing benchmarks mainly evaluate medium-horizon exploration and rarely test whether agents can sustain long, dependency-heavy research processes. We introduce Mr.LHDR (Multimodal real-world Long-Horizon Deep Research), a benchmark for evaluating real-world deep research over long, irreducible chains of interdependent evidence across eight categories. Each question is constructed from a hidden Node-Relation graph and requires an average of 12.1 necessary intermediate conclusions with a mean dependency depth of 10.4 before reaching a short, unique, and verifiable answer. Questions incorporate multimodal evi
קרא במקור המקורי