יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

ExplorationBench: מדידת חקירה במערכות AI

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds
ExplorationBench הוא כלי למדידת יכולת החקירה של מערכות AI. הוא מאפשר לבדוק יכולתן של מערכות AI לגלות ידע חדש וליישמו בסביבות לא מוכרות. הכלי כולל שני סנדבוקסים: AlienCode ו-AlienLogic, שמספקים משוב סביבתי וכלים לחקירה.
תקציר מקורי באנגליתarXiv:2609.30199v2 Announce Type: replace-cross Abstract: Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However, evaluating this ability is difficult: (1) how to verify whether a genuinely new hypothesis holds, and (2) how to determine whether a system has discovered it through exploration or merely recalled related knowledge from pre-training data. To this end, we introduce ExplorationBench, which turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds: their rules are executable, so every answer can be checked exactly, and they conflict with familiar knowledge, so recall alone cannot s
קרא במקור המקורי