כתבה
arXiv cs.AI ·
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite
תקציר מקורי באנגליתarXiv:2610.02826v1 Announce Type: new Abstract: Successful trajectories on difficult tasks provide valuable supervision for model improvement, but specialized harnesses introduce interventions that may be unavailable during deployment. We propose Recursive Self-Rewrite (RSR), a framework that uses one base model, Qwen-3.8-27B, to discover successful solutions under diverse harnesses and reconstruct them as training trajectories under a general harness. A planner extracts procedures into runbooks, a critic screens for verifier and solution leakage and guides recursive revision, and an executor follows qualified runbooks in fresh sandboxes. Across approximately 3K self-curated terminal tasks, three harnesses jointly solve 759 tasks, 34.3% more than the strongest individual harness in the rec
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית