כתבה
arXiv cs.AI ·
SeRV: תקיפה-מאורגנת של רזולוציית וקטורים עבור ייצור ASL
SeRV: Semantic-Aligned Residual Vector Quantization for American Sign Language Generation
SeRV היא טכנולוגיה חדשה לייצור ASL שמשתמשת בשיטת רזולוציית וקטורים מאורגנת. היא מסוגלת לייצר תנועות ASL מדויקות ומובנות, כולל תנועות 3D.
תקציר מקורי באנגליתarXiv:2609.05742v1 Announce Type: cross Abstract: American Sign Language (ASL) generation remains challenging due to limited paired text-ASL motion data and the difficulty of learning motion representations both precise for reconstruction and predictable from linguistic input. Existing methods rely on motion tokenizers optimized for reconstruction, without explicit semantic supervision from paired text. As a result, the learned tokens remain limited in supporting semantically consistent and fine-grained ASL motion generation. To address this limitation, we propose SeRV (Semantic-Aligned Residual Vector Quantization), a semantic-aligned RVQ tokenizer for ASL generation. SeRV learns a semantically structured residual token space by combining sentence-level motion-text alignment with token-le
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית