כתבה
arXiv cs.LG ·
TraceML: What Auto-Research Agents Miss in Long-Horizon ML Development
תקציר מקורי באנגליתarXiv:2608.26086v3 Announce Type: replace Abstract: Auto-research agents now run machine-learning development unattended for hours, revising data pipelines, models, and validation from their own feedback, yet on most competitions they still finish below strong human competitors. Outcome-based benchmarks record this gap but not its cause, because they grade the final submission and discard the development process behind it. We introduce TraceML, which pairs human and agent work on the same competitions under one version-level schema: 4{,}465 human Kaggle trajectories across 134 competitions, seven of which are also worked by two agent scaffolds, giving 430 paired human and 207 agent trajectories. Every code version carries its score, its timestamp, and labels for the action taken, its inten
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית