כתבה
arXiv cs.AI ·
פרוטוקול CaRE להערכת מודלים של שפה מבוזרים
CaRE Compute-aware Remasking Evaluation Protocol for Masked Diffusion Language Models
פרוטוקול CaRE הוא כלי להערכת מודלים של שפה מבוזרים. הוא בודק את הביצועים של מודלים כמו LLaDA-8B-Base ו-Dream-7B-Base. הפרוטוקול מאפשר השוואה בין מודלים שונים ומחשב את הביצועים שלהם.
תקציר מקורי באנגליתarXiv:2607.24763v1 Announce Type: new Abstract: Masked diffusion language models (MDLMs) are advancing rapidly, yet the evaluation standards needed to reliably interpret their progress have not kept pace. Despite MDLMs becoming competitive with autoregressive language models, seven recent remasking papers evaluate under incompatible settings, varying nominal step counts, metrics, and sampling temperatures without jointly controlling these factors, rendering their strategy rankings largely incomparable and leaving open whether reported gains reflect algorithmic improvements or evaluation artifacts. We present CaRE, a compute-aware evaluation framework that audits MDLM remasking strategies by standardizing actual number of function evaluations (NFE), enforcing multi-metric reporting, and exp
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית