כתבה
arXiv cs.LG ·
למידת ריפוד עם קבוצות פעולה מותאמות: אפליקציה למלצות רצפתיות
Reinforcement Learning with Conformal Action Sets: An Application to Sequential Recommendation
אפליקציה ללמידת ריפוד עם קבוצות פעולה מותאמות למלצות רצפתיות. המאמר עוסק בפיתוח שיטה חדשה ללמידת ריפוד, המשתמשת בקבוצות פעולה מותאמות לשיפור תוצאות המלצות רצפתיות. השיטה, הקרויה RLCP, משתמשת בביקורת סקורים ובעריכה מתקדמת של קבוצות פעולה. המאמר כולל תיאור של השיטה, תיאור של המבחן, ותוצאות המבחן. התוצאות המבחן מראות כי השיטה החדשה משפרת את תוצאות המלצות רצפתיות בהשוואה לשיטות למידת ריפוד קיימות.
תקציר מקורי באנגליתarXiv:2610.08743v1 Announce Type: new Abstract: Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session. We propose Reinforcement Learning with Calibrated Pruning (RLCP), which adapts the retained action set using critic scores and an online threshold. The threshold is updated from binary feedback indicating whether the set contains an action in a proxy target. We prove a deterministic bound on the observed proxy miss rate along adaptive trajectories. To quantify the effect of pruning on reward, we derive an exact decomposition of value loss into filtering and selection losses. Under explicit proxy and critic approximation conditions, this decomposition yields a finite session reward bound that also accounts for imperf
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית