כתבה
arXiv cs.LG ·
AhaBench: האם Agent לומד מהתקופה הקודמת? בסיס תצוגה ללמידה רציפה באורך-זמן
AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning
בסיס תצוגה ללמידה רציפה באורך-זמן, המבחן את יכולת ה-Agent ללמוד מהתקופה הקודמת. הבסיס כולל שלושה חלקים: Aha-Puzzle, Aha-Euler ו-Aha-Vending. הבסיס נועד לבחון את יכולת ה-Agent להשתפר באופן רציף וללמוד מהתקופה הקודמת.
תקציר מקורי באנגליתarXiv:2609.05435v1 Announce Type: new Abstract: Modern language agents are expected to operate over long horizons: they ask follow-up questions, reuse worked examples, handle tool feedback, and adapt to delayed consequences. Most evaluations still reset the agent after a prompt or score only the final state of one trajectory. AhaBench asks a more operational question: when a fixed model receives useful experience, does its later behavior improve under a related evaluation condition where the obvious support has been removed, changed, or delayed? The suite contains three components. Aha-Puzzle tests no-hint exploration after solved hidden-state puzzles; Aha-Euler turns Project-Euler-style mathematical ideas into generated taught/held-out tasks with exact validators; and Aha-Vending, an open
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית