כתבה
arXiv cs.AI ·
TraceML: עד כמה סוכלו סוכני-מחקר אוטומטיים בפיתוח ML לאורך זמן
TraceML: What Auto-Research Agents Miss in Long-Horizon ML Development
סוכני-מחקר אוטומטיים נכשלים בפיתוח ML לאורך זמן. חוקרים הציגו חידוש בשם TraceML, שמציג את הפעילות של סוכני-מחקר ואנשי מחקר יחד. החידוש נועד לצפות את הפעילות של סוכני-מחקר ולסייע לשפר את תוצאותיהם.
תקציר מקורי באנגליתarXiv:2608.26086v3 Announce Type: replace-cross Abstract: Auto-research agents now run machine-learning development unattended for hours, revising data pipelines, models, and validation from their own feedback, yet on most competitions they still finish below strong human competitors. Outcome-based benchmarks record this gap but not its cause, because they grade the final submission and discard the development process behind it. We introduce TraceML, which pairs human and agent work on the same competitions under one version-level schema: 4{,}465 human Kaggle trajectories across 134 competitions, seven of which are also worked by two agent scaffolds, giving 430 paired human and 207 agent trajectories. Every code version carries its score, its timestamp, and labels for the action taken, its
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית