יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

החלק הקשה: במבחן ובעזרה - כיצד חברות פיתחו חוקרים חדשים לבדיקת זיכרון והצגה

The Hard Part Comes After Search: Benchmarking Web Agents on Synthesizing, Organizing, and Displaying Knowledge
במבחן חדש, חוקרים בדקו את יכולתם של חברות לפתח חוקרים חדשים לבדיקת זיכרון והצגה. המחקר, שפורסם בארכיון arXiv, חשף חולשות בזיכרון ובהצגה של חברות פיתחו חדשים.
תקציר מקורי באנגליתarXiv:2609.30604v1 Announce Type: new Abstract: Existing computer-use agent benchmarks do not fully evaluate agents acting as assistants. A useful assistant retrieves information across complex, multi-step workflows, synthesizes it into artifacts (documents, presentations, spreadsheets), and navigates program interfaces to produce a coherent final product. Such workflows demand reasoning and synthesis, decomposition of complex tasks, as well as visual and spatial understanding. To study agents on workflows like these, we introduce KNOWS, a benchmark of open-ended, complex, browser-based tasks that jointly evaluate these capabilities, with each task culminating in a produced artifact. To write tasks, we develop a task design rubric and a protocol for ensuring that tasks meet the requirement
קרא במקור המקורי