יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הנהיגה בלי לפגוע: פעולות התערבות מומחים למודלי דיפוזיה דיסקרטית לשפה

Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models
מחקר חדש מציע פעולות התערבות מומחים למודלי דיפוזיה דיסקרטית לשפה, המשפרות את היחס האיכות-שליטה. המחקר נעשה על מודלי DLMs (124M-8B פרמטרים) ומציע שיטה אדפטיבית להפעלת התערבות, המתמקדת באזורים שבהם כל תכונה נוצרת. השיטה מוכיחה עצמה כיעילה יותר משיטות התערבות יומוריות ומגבילות.
תקציר מקורי באנגליתarXiv:2605.10971v2 Announce Type: replace Abstract: Discrete diffusion language models (DLMs) generate text by iteratively denoising all positions in parallel, offering an alternative to autoregressive models. Controlled generation methods for DLMs, imported from autoregressive models, apply uniform intervention at every denoising step. We show this uniform schedule is inefficient and degrades quality, and the damage compounds when multiple attributes are steered jointly. To diagnose the failure, we train sparse autoencoders on four DLMs (124M-8B parameters) and find that different attributes commit on distinct schedules, varying in timing, sharpness, and magnitude. For instance, topic commits within the first 2% of denoising on MDLM, whereas sentiment emerges gradually over 20% of the pro
קרא במקור המקורי