יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

אופטימיזציה של קוד על ידי למידה עצמית

Reinforcement Learning for Code Optimization
אופטימיזציה של קוד על ידי למידה עצמית. המחברים פיתחו שיטה לאופטימיזציה של קוד שמשתמשת בלמידה עצמית. השיטה כוללת שלושה שלבים: (1) הקמת סביבת ניסויים, (2) חישוב הפרפורמנס של הקוד, ו(3) אימון המודל. המחברים הציגו תוצאות של קוד שאופטימזציה על ידי שיטתם.
תקציר מקורי באנגליתarXiv:2607.25970v2 Announce Type: replace Abstract: RL for code correctness is now established: have the model generate a program, run it against hidden test cases, and reward solutions that pass. Extending this to code optimization seems straightforward: just add execution time to the reward. But in practice, once timing drives the reward, small problems in measurement noise, reward sparsity, or GRPO instability overwhelm the signal and make RL fail: generated solutions are barely faster, and more of them can fail. We make execution time learnable through three stages: (1) how code is tested, by building DMC-Optim with large optimization tests and a calibrated sandbox; (2) how speed is turned into reward, by composing correctness and speed in the RL environment and using an offline simula
קרא במקור המקורי