כתבה
arXiv cs.AI ·
EngramBench: A Capability-Grounded Benchmark for Skill-Evolution Harnesses
תקציר מקורי באנגליתarXiv:2609.39284v1 Announce Type: cross Abstract: While large language models have achieved remarkable success in isolated code generation, authentic software engineering requires sustained reasoning, complex state management, and continuous cross-domain abstraction. However, current evaluations of skill evolution in autonomous agents suffer from a critical identifiability problem: they structurally confound genuine capability abstraction with rote solution leakage (i.e., copying highly similar code from historical training data). To resolve this, we introduce EngramBench, a rigorous, capability-grounded benchmark governed by the strict axiom of capability overlap without solution overlap. Comprising 30 diverse learning tasks and 13 unseen transfer tasks, EngramBench challenges agents to n
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית