יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

בלמן פוגש ליאפונוב: למידת חיזוק לא מ�ורטת דרך התמצאות בכאוס

Bellman Meets Lyapunov: Unsupervised Reinforcement Learning via Mastering Chaos
חוקרים מציגים שיטה חדשה ללמידת חיזוק לא מ�ורטת, המאפשרת לסוכנים ללמוד התנהגויות בסיסיות מבלי להידרש לאותות רווח מובנים. השיטה, הנקראת Forward CIP, מבוססת על עקרונות דינמיים בלבד ואינה דורשת ידע תחומי מוקדם.
תקציר מקורי באנגליתarXiv:2610.02012v1 Announce Type: new Abstract: Reinforcement learning (RL) is a powerful paradigm for training agents, yet its success rests on domain expertise of human engineers who design informative reward signals for every new task. Unsupervised RL aims to reduce this engineering with intrinsic motivation (IM): reward signals that emerge from the agent environment interaction itself. Existing IM objectives, however, involve the selection of information variables, which re-introduces domain expertise the field has sought to eliminate. We introduce Forward CIP (F-CIP), an RL-native formulation of the Controllable Information Production (CIP) objective, which is defined by the system's dynamics alone and requires no such selection. We prove that F-CIP is compatible with RL and demonstra
קרא במקור המקורי