כתבה
arXiv cs.AI ·
Not Too Hard, Not Too Easy: Learning from Intermediate States for LLM Structured Reasoning
תקציר מקורי באנגליתarXiv:2609.33149v2 Announce Type: replace Abstract: A common principle of effective learning is to practice material that is neither already mastered nor too difficult to permit progress. We ask how to apply this principle to structured reasoning tasks such as Sudoku and maze solving. In these tasks, a model can repeatedly revise an incomplete or incorrect candidate solution until it satisfies the problem's constraints. The intermediate candidate solutions along this trajectory provide natural training examples: some are already solved, some cannot yet be repaired by the model, and others lie at its current frontier of achievable progress. We therefore investigate whether pretrained language models can learn to revise such states and whether training on states at this frontier improves rea
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית