כתבה
arXiv cs.AI ·
LineupRL: למידת חיזוק עם תגמולים מאומתים
LineupRL: Verifiable Reinforcement Learning for Time Series Captioning via Caption-to-Series Identification
LineupRL הוא פרויקט שמציג שיטה חדשה ללמידת חיזוק עם תגמולים מאומתים. השיטה משתמשת במודל לשוני גדול כדי ליצור תגמולים עבור כותרות סדרות זמן. LineupRL מראה תוצאות טובות יותר משיטות אחרות בתחום.
תקציר מקורי באנגליתarXiv:2610.01800v1 Announce Type: new Abstract: Time series captioning is a fundamental step in time series understanding and can also serve as the bridge between signal and natural language. Supervised fine-tuning (SFT) relies on a larger model's captions and cannot exceed their quality. Reinforcement learning (RL) can, but its rewards were designed for other modalities and other tasks, and they transfer poorly to open-ended generation in the time series domain. We address this by proposing LineupRL, a reinforcement learning with verifiable rewards (RLVR) pipeline whose reward is caption-to-series identification. The reward model is a frozen large language model (LLM) verifier that reads the generated caption and the candidate time series as raw values, never the chart, and must pick the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית