כתבה
arXiv cs.AI ·
Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models
תקציר מקורי באנגליתarXiv:2609.09957v1 Announce Type: cross Abstract: Modern large language models (LLMs) can translate natural-language descriptions into operations research (OR) formulations. Post-training techniques including reinforcement learning and on-policy self-distillation have further improved this capability. However, three limitations remain in training LLMs for OR formulations. First, training commonly relies on synthetic formulations validated by human experts or stronger models, constraining scalable supervision. Second, credit assignment is either coarse or costly: outcome rewards score an entire trajectory without locating the responsible modeling decision, whereas process-level supervision requires an additional evaluator. Third, privileged self-distillation can induce style mismatch by usi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית