כתבה
arXiv cs.LG ·
Code-to-Harness: Distilling Black-Box Optimizers from Self-Play
תקציר מקורי באנגליתarXiv:2609.09468v2 Announce Type: replace Abstract: Can an agent learn a numerical search strategy through executable practice and then transfer that strategy as text? We study low-budget black-box optimization, where unaided language models remain well below strong classical optimizers. During development, an agent repeatedly writes and evaluates optimizer programs. It then distills the resulting program and practice record once into a 197-word primary Harness A, which is frozen before evaluation. Harness A reduces Gemini Flash regret by 48\% in an independent $N=30$ study ($p<.001$), enters the GP-BO performance range on the practice family, and lowers mean regret on all three held-out BBOB landscapes. The same text improves every tested Gemini executor and transfers to Claude Sonnet, re
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית