כתבה
arXiv cs.LG ·
פיתוח אגנטים אוטומטיים על ידי דגמי עולם
Scaling Automatic Research Agents via World Models
אגנטים אוטומטיים נמדדים על ידי דגמי עולם. פיתוח זה מאפשר גידול באפקטיביות של 3-4 פעמים, ומעבר פורמטיבי של 4B ו-9B ל-48B ו-120B.
תקציר מקורי באנגליתarXiv:2608.12564v3 Announce Type: replace Abstract: Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch) agents bring this goal within reach, as modern LLMs show the capability to independently implement solutions and learn from the execution outcomes. Behind these gains, post-training (especially RL) plays a central role. In this paper, we identify a fundamental tension when scaling RL for these agents: the two components of every AutoResearch trajectory (agent generation and environment execution) scale in very different manners, since all generation shares compute through batching, while each execution occupies its exclusive sandbox and real machine time. As a result, the environment execution dominates the training cost and becomes
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית