כתבה
arXiv cs.AI ·
פיתוח פונקציות שכר ללמידת רכיבה באופן חסכוני
Specifying Reward Functions for RL Without Environment Sampling
אפשרות חדשה לפיתוח פונקציות שכר ללמידת רכיבה, ללא צורך בגישה למערכת. פיתוח זה עושה שימוש ב-LLM ובאימות מבוסס על ניסוחי ניהול. המאמר עוסק בפיתוח פונקציות שכר עבור שלושה תחומים: ניהול פנדמיה, ניהול סוכרת וניהול רכב נפח.
תקציר מקורי באנגליתarXiv:2609.15544v1 Announce Type: cross Abstract: Enabling human stakeholders to specify reward functions that lead to their desired outcomes is a key challenge in deploying reinforcement learning agents. Preference-based methods such as online RLHF can reduce the burden of manual reward design, but they require repeatedly training policies, sampling trajectories from the real world, and eliciting feedback, making them impractical in settings where environment interaction is computationally expensive or unsafe. We introduce Experience-Free Autonomous Reward Specification (EARS), a method for learning reward functions from preferences without environment interaction. Our approach uses a structured LLM-mediated process to construct a small set of expressive reward features from a task descri
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית