יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

תרגיל תפיסה חזותי-קולי בסצנות פארטי

Audio-Visual Turn-taking Prediction in Cocktail Party Scenarios
מודלי תפיסה חזותי-קולי לתפיסת תורים נכשלים בשיחות עם רקע רעשי.
תקציר מקורי באנגליתarXiv:2609.17056v3 Announce Type: replace-cross Abstract: Current predictive turn-taking models (PTTMs) achieve strong performance on benchmarks with controlled acoustic conditions and clean audio signals. Their generalisation to conversations with overlapping speech and background interference remains underexplored. In this research, we evaluate audio-visual PTTMs trained with clean data on a challenging cocktail-party testbed derived from the AVCocktail dataset, and analyse their adaptation behaviour to this new domain. Experimental results show consistent performance degradation across audio and visual modalities under noisy conditions, with up to 38% relative drop in weighted F1. Fine-tuning on the new domain improves robustness, but gains vary across modalities and depend on the size
קרא במקור המקורי