יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

עיצוב מדיניות מטא בטוח עם בקרת סיכון

Safe Meta-Policy Design with Risk Control
חוקרים פיתחו שיטה לעיצוב מדיניות מטא בטוחה עם בקרת סיכון, המאפשרת לתכנ� עדכוני מדיניות לפני שהם מוכשרים, תוך מאזנת בין יתרונות השיפור לסיכון של ירידה בביצועים.
תקציר מקורי באנגליתarXiv:2610.10393v1 Announce Type: cross Abstract: Models can be retrained as new data arrive, but deploying every new version risks replacing a good policy with a worse one. We study how to plan policy updates (i.e., meta-policy) before future candidates are trained, balancing the benefits of improvement against the risk of performance regression. Our offline meta-policy maximizes expected cumulative value subject to a budget on the expected number of updates that perform worse than the policies they replace. We estimate the value and risk of possible switches from historical learning trajectories, represent an update schedule as a path in a directed acyclic graph, and select a schedule using dynamic programming. A leading-order analysis identifies the signal-to-noise ratio of policy impro
קרא במקור המקורי