יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מחפש לסיגנל: עדכון אונליין לאחר הכשרה בעיצוב היוריסטי האוטומטי

From Search to Signal: Online Post-Training in Automatic Heuristic Design
מחפש לסיגנל: עדכון אונליין לאחר הכשרה בעיצוב היוריסטי האוטומטי. המחקר מציג פתרון לבעיה של עדכון רכיבי ה-LLM בעיצוב היוריסטי האוטומטי. הפתרון משתמש בשיטת RLVR ומציע פתרון לבעיה של עדכון רכיבי ה-LLM בעיצוב היוריסטי האוטומטי.
תקציר מקורי באנגליתarXiv:2609.39383v1 Announce Type: cross Abstract: Large language model (LLM)-based automatic heuristic design (AHD) iteratively proposes and refines heuristics, pairing design rationales with executable code. Task-specific evaluators assess programs; execution outcomes and performance scores guide search. Many AHD systems keep the generator frozen; EvoTune and Co-Evolution of Algorithms and Language Model (CALM) instead update it from evaluated candidates. When such outcomes drive reinforcement learning with verifiable rewards (RLVR), they create a search-coupled loop: the evaluated candidate stream supplies both search-state updates and training signals for the model that generates future candidates. Yet validity and performance do not uniquely determine useful model updates; converting t
קרא במקור המקורי