כתבה
arXiv cs.LG ·
CARM: מסיכת תגובה המודעת לביטול
CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning
CARM היא שיטה חדשה לשליטה בתגובות מחוץ למדיניות בלמידת לשונית חיזוקית. היא משתמשת במסיכת תגובה המודעת לביטול כדי לשפר את היכולת של מודלים כמו LLaMA. השיטה הוכחה כיעילה בניסויים.
תקציר מקורי באנגליתarXiv:2610.02039v1 Announce Type: new Abstract: Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make sampled responses off-policy. Sequence-level masking addresses this mismatch by deciding whether an entire response should contribute to optimization. A common masking rule uses the length-normalized geometric mean of sampled token probability ratios. Its signed log-ratios can cancel across positions, concealing substantial bidirectional policy drift. We propose \emph{Cancellation-Aware Response Masking} (CARM), a sequence-level mask that takes th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית