יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

חקירה והסברה: פענוח סיכונים רציונלי בדיפוזיה של דגלי תקשורת

Exploring More, Reasoning Better: Stepwise Risk-Sensitive GRPO for Diffusion Language Models
מאמר חדש: פיתוח דגלי תקשורת לשפה של המודלים, עם פענוח סיכונים רציונלי. המאמר מציע שיטה חדשה לשיפור דגלי תקשורת, המבוססת על פענוח סיכונים רציונלי. השיטה, הנקראת StepRS-GRPO, משתמשת בטכניקה של GRPO (Group Advantage Transformation) כדי לשפר את הדגלי תקשורת. המאמר כולל ניסויים שונים, המציגים את יעילות השיטה. המאמר גם כולל דיון בקשר של השיטה לדגלי תקשורת אחרים.
תקציר מקורי באנגליתarXiv:2610.00661v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) generate text by denoising a sequence or successive blocks, allowing several tokens to be revealed in parallel. Reinforcement learning with verifiable rewards (RLVR) reuses terminal feedback across these decisions, even as their conditioning context changes. We propose stepwise risk-sensitive GRPO (StepRS-GRPO), which varies the risk coefficient of the group-advantage transformation across denoising states while retaining the underlying trainer. For binary rewards, we show that this transformation is exactly a prompt- and state-dependent rescaling of centered outcome advantages. A capability-based calibration suggests a coefficient scale, while endpoint and interpolation ablations guide schedule selecti
קרא במקור המקורי