כתבה
arXiv cs.AI ·
Open-UniMo: איחוד תנועה-שפה
Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World
Open-UniMo הוא מודל שפה-תנועה מאוחד המאפשר ייצוג משותף של שפה ותנועה. המודל משתמש ב-Qwen ומאפשר יצירת תנועה מתוך טקסט והבנת תנועה לטקסט.
תקציר מקורי באנגליתarXiv:2609.14615v1 Announce Type: cross Abstract: Unified motion generation and understanding is crucial for embodied AI systems that can both synthesize and interpret human actions in open-world environments. Existing motion-language models often treat motion as an auxiliary modality of a language model, leading to text-dominated representations and limited cross-modal interaction. Moreover, the next-token prediction paradigm is not naturally suited to long motion sequences, where autoregressive generation may accumulate prediction errors. To address these challenges, we propose Open-UniMo, a unified Large Motion-Language Model (LMLM) trained on million-scale open-world motion-language data. Open-UniMo promotes modality parity by extending Qwen's vocabulary of about 150K text tokens with
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית