יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

SemanTok: טוקנים סמנטיים ניתנים לחיזוי לייצור וידאו אוטורג'סיבי

SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
SemanTok הוא טוקניזר וידאו רכיב-סמנטי שמספק טוקנים סמנטיים ניתנים לחיזוי. הוא משתמש ב-DINO כמקור לטוקנים ומספק רכיבים קלים שמחזירים את הטוקנים מכל פרפקס של טוקן. SemanTok משפר את התאמת הסמנטיקה ואת האמינות של הווידאו בכל גודל של AR. הוא עובד טוב בשני תחומים: ריפוי וייצור. הטוקנים הקצרים שלו זולים יותר לחיזוי ומספקים אמינות טובה יותר.
תקציר מקורי באנגליתarXiv:2610.00686v1 Announce Type: cross Abstract: Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, coarse-to-fine tokenizers yield exactly that: the first coarse tokens carry the clip's global semantics while later tokens further specify details. Existing flexible tokenizers only apply a representation-alignment (REPA) loss on early decoder hidden states, a target the decoder can partly meet from its noised input instead. We introduce SemanTok, a flexible video tokenizer that feeds frozen DINO features into its encoder and adds lightweight heads that reconstruct t
קרא במקור המקורי