יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

FloorSAV: הסברת תצורה אודיו-ויזואלית-2D עם פלורמפ 2D ל-LLMs AV

FloorSAV: Elucidating Spatial Audio-Visual Context with 2D Floormap for AV-LLMs
FloorSAV הוא פלטפורמה שמסבירה תצורה אודיו-ויזואלית-2D עם פלורמפ 2D דינאמי. היא משלבת עובדות חזותיות, אודיותיות וגאומטריות כדי לשפר את היכולת של LLMs להבין תצורה ספציפית. FloorSAV נבחן ב-SAVED-Bench ו-SAVVY-Bench והוכיח יכולת טובה בעבודה עם תצורה ספציפית.
תקציר מקורי באנגליתarXiv:2610.11310v1 Announce Type: cross Abstract: While 3D spatial reasoning in dynamic egocentric environments is crucial for embodied intelligence, audio-visual large language models (AV-LLMs) lack explicit mechanisms to process and internalize global geometry directly from raw sensory streams. Existing approaches either require costly fine-tuning or underutilize the model's cross-modal reasoning capacities. In this paper, we propose FloorSAV, a novel framework that explicitly grounds spatial audio-visual context by rendering a dynamic 2D floormap. By integrating 3D point clouds, camera trajectories, spatial audio cues, and semantically grounded object landmarks, we inject this floormap into the AV-LLM as a synchronized stream with an egocentric video. AV-LLMs utilize their multi-modal c
קרא במקור המקורי