כתבה
arXiv cs.AI ·
Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis
תקציר מקורי באנגליתarXiv:2610.03548v1 Announce Type: new Abstract: Generating progressively harder reasoning problems requires synthesis procedures that adapt as the task distribution evolves. Existing task-level recursion reuses generated problems as seeds but leaves the construction harness unchanged. We present task-harness co-evolution, a framework for recursive harness self-improvement (RSI) in reasoning-data synthesis. Online self-improvement converts intermediate solver failures into reusable skills during generation. Post-task self-improvement revises skills, prompts, and workflows after each batch, adopting candidates only when they generate harder valid tasks within a bounded cost increase. Model weights and verification criteria remain fixed. Across mathematics, coding, and science, mean solver ac
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית