כתבה
arXiv cs.AI ·
EurekaBench: מדד ליכולת האגנטים לגלות תגליות חדשות
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
EurekaBench היא תקן חדש שמבדוק את יכולת האגנטים לגלות תגליות חדשות. התקן כולל 26 תרגילים ארוכי-זמן בתחומים שונים, כולל פיזיקה, כימיה וביולוגיה.
תקציר מקורי באנגליתarXiv:2610.00492v1 Announce Type: cross Abstract: When Isaac Newton discovered the law of gravitation, he did so through an iterative process of analyzing observed data such as planetary patterns, finding the underlying mechanisms by describing patterns in mathematical equations, and refining his theory against the Moon's orbit, revealing the startling insight that the same force governs both falling apples and orbiting planets. Would it be possible for AI agents to make similar discoveries? To measure this ability, we introduce EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. We evaluate these mechanisms by the scientific insights that can be derived from them. EurekaBench contains an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית