כתבה
arXiv cs.LG ·
From Search to Signal: Online Post-Training in Automatic Heuristic Design
תקציר מקורי באנגליתarXiv:2609.39383v1 Announce Type: new Abstract: Large language model (LLM)-based automatic heuristic design (AHD) iteratively proposes and refines heuristics, pairing design rationales with executable code. Task-specific evaluators assess programs; execution outcomes and performance scores guide search. Many AHD systems keep the generator frozen; EvoTune and Co-Evolution of Algorithms and Language Model (CALM) instead update it from evaluated candidates. When such outcomes drive reinforcement learning with verifiable rewards (RLVR), they create a search-coupled loop: the evaluated candidate stream supplies both search-state updates and training signals for the model that generates future candidates. Yet validity and performance do not uniquely determine useful model updates; converting the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית