כתבה
arXiv cs.CL ·
דינויזינג של ייצוגים היררכיים: דיפוזיה רצפתית משותפת ללמידת מודלי שפה
Denoising Hierarchical Representations: Joint Continuous Diffusion for Language Modeling
אנו מציגים פרקטיקה חדשה לשיפור של מודלי למידת שפה רצפתיים, המשתמשת בייצוגים היררכיים. הפרקטיקה, המכונה H-CDLM, משתמשת בדיפוזיה רצפתית משותפת לשני רמות ייצוג: רמת הטקסט עצמו ורמת קבוצות טקסט שנוצרו על ידי קלסטריזציה של ייצוגי טקסט מוכנים. הפרקטיקה מציעה קבוצת ניסוחים כלליים שמאפשרים סיימרים ומתכונים שונים לכל רמה. הפרקטיקה נבחנה על CoBit והציגה תוצאות טובות יותר. הפרקטיקה גם נבחנה על FLM והציגה תוצאות טובות יותר.
תקציר מקורי באנגליתarXiv:2610.08738v1 Announce Type: new Abstract: Diffusion Language Models (DLMs) hold the promise of order-agnostic, parallel text generation. Recently, continuous diffusion and flow matching models have seen substantial gains, driven by carefully crafted token representations and diffusion/flow spaces. In this work, we introduce Hierarchical Continuous Diffusion Language Models (H-CDLMs), a simple framework that further improves continuous DLMs with minimal compute and parameter overhead. Drawing on the discrete DLM and continuous image diffusion literature on joint diffusion, we diffuse multiple modalities in parallel. These modalities represent tokens at different semantic granularities: in our instantiation, the tokens themselves and coarser clusters obtained by clustering pretrained t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית