כתבה
arXiv cs.LG ·
למידה בטוחה תחת דינמיקה אירוציונלית: להתבקש לעזרה
Safe Learning Under Irreversible Dynamics via Asking for Help
במאמר זה, החוקרים מציעים פתרון ללמידה בטוחה במערכות שבהן טעויות אינן ניתנות לתיקון. הם מציעים ללמידה להתבקש לעזרה ממורה ולהעביר ידע בין מצבים דומים. התוצאות המדעיות הן חשובות, והן יכולות לשפר את היכולת של הלמידה להשיג פרסים גבוהים באזורים חדשים.
תקציר מקורי באנגליתarXiv:2502.14043v4 Announce Type: replace Abstract: Most learning algorithms with formal regret guarantees essentially rely on trying all possible behaviors, which is problematic when some errors cannot be recovered from. Instead, we allow the learning agent to ask for help from a mentor and to transfer knowledge between similar states. We show that this combination enables the agent to learn both safely and effectively. Under standard online learning assumptions, we provide an algorithm whose regret and number of mentor queries are both sublinear in the time horizon for Markov decision processes with irreversible dynamics and infinite state spaces. Our proof involves a sequence of three reductions, making our result more general than a single algorithm. Conceptually, our result may be the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית