יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

AgentIdeaBench: נקודת דיווח להערכת רעיונות מדעיים בעידן האג'נט

AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era
נקודת דיווח להערכת רעיונות מדעיים בעידן האג'נט. AgentIdeaBench מבחנת את יכולת המודלים להציע רעיונות חדשים ומבחינה בין המודלים השונים.
תקציר מקורי באנגליתarXiv:2609.07611v1 Announce Type: new Abstract: Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and autonomous AI scientists depend on it. Existing evaluations largely assess it by asking models to generate ideas from a static, curated set of reference papers. That passive setup departs from the retrieval-and-reasoning workflow of modern AI scientists, and it becomes less discriminative as models improve. We introduce AgentIdeaBench, a multidisciplinary benchmark that evaluates scientific ideation under two matched settings, static observation and active exploration. We report matched Static-Active evaluations for 33 LLMs across 40 densely scored subfields spanning five disciplines, using a multidimensional, literature-verified scori
קרא במקור המקורי