כתבה
arXiv cs.AI ·
מודלים רציפים גדולים לשפה
Large Language Continuous Diffusion Models
מודל Sigma הוא מודל שפה רציף גדול שמשתמש במסלולים סמויים ניתנים לכיוון. הוא מאיץ אימון עם משקולות מוכנות מראש ממודלים אוטורגרסיביים. Sigma משיג ביצועים תחרותיים עם מודלים בדיסקרטיים בבדיקות מתמטיקה וקוד.
תקציר מקורי באנגליתarXiv:2610.02665v1 Announce Type: cross Abstract: Despite the success of discrete diffusion language models (dLMs) for fast parallel decoding, their non-smooth, high-dimensional space hinders trajectory steering for reasoning and inference acceleration. To overcome this, we present Sigma, the first large-scale (3B/8B) continuous dLM built on steerable, low-dimensional ODE/SDE latent trajectories. Trained blockwise via likelihood optimization, Sigma jointly denoises Gaussian-corrupted token embeddings while learning an optimal embedding geometry. To accelerate training, Sigma leverages pre-trained weights from autoregressive (AR) models for warm-starting. During inference, we identify classifier-free guidance and score temperature as essential for high-fidelity reasoning and coding. Across
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית