יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

אפשרות לאפיין תנועות 3D של מספר אנשים רק באמצעות צליל

Sound-based Multi-Person 3D Pose Estimation
במאמר זה נציגים את SoundMHPE, פרקטיקה חדשה לאפיין תנועות 3D של מספר אנשים רק באמצעות צליל. הפרקטיקה כוללת שני חלקים: Acoustic Multi-scale Encoder ו-Temporal Pose Decoder. ה-Acoustic Multi-scale Encoder יוצר תמונה כוללת של תנועות 3D של מספר אנשים, וה-Temporal Pose Decoder פועל לפי תפקוד של רשת תקשורת, ומאפשר לאפיין תנועות 3D של מספר אנשים. המאמר כולל גם תיאור של AMP, סט הנתונים החדש, ובו 432,000 תמונות של תנועות 3D של מספר אנשים.
תקציר מקורי באנגליתarXiv:2609.04902v1 Announce Type: cross Abstract: Can we recover the 3D poses of multiple people using only sound? This paper presents the first attempt to estimate multi-person 3D poses solely from acoustic signals. Estimating the poses of multiple individuals using acoustic signals is inherently challenging due to the superposition of motion-dependent signal variations. Unlike single-person scenarios, the presence of multiple subjects leads to overlapping acoustic signatures, making it difficult to attribute specific signal changes to an individual's pose. Furthermore, the complexity is compounded by inter-person reflections, which introduce intricate propagation delays that obscure the temporal motion-acoustic relationship. To address these issues, we propose SoundMHPE (Sound-based Mult
קרא במקור המקורי