כתבה
arXiv cs.AI ·
פחות יכול להיות יותר: מהו היבטים של דיבור שמניעים גילוי סוף-תור
Less can be More: What Aspects of Speech Drive End-of-Turn Detection
זיהוי סוף-תור בבינה מלאכותית קונברסית מסתמך על אותות אקוסטיים ופרוזודיים. מחקר זה מראה כי שילוב אקוסטי-פרוזודי משיג את האיזון הטוב ביותר בין דיוק ועיכוב, עם F1 של 0.93 ו-7.8% התראות שווא. ניתוח מרחב המאפיינים מאשר כי מאפיינים פרוזודיים הם בעלי ההפרדה החזקה ביותר, בעוד ייצוגי טקסט משותפים במידה ניכרת.
תקציר מקורי באנגליתarXiv:2609.11066v1 Announce Type: cross Abstract: In conversational AI, detecting when a speaker has finished talking is crucial for natural turn taking. While recent work incorporates semantics, the relative contribution of different modalities remains unclear. We present a controlled ablation of acoustic, prosodic, and semantic signals for streaming end of turn detection using a lightweight trimodal classifier. Under identical training conditions, the acoustic prosodic combination achieves the best balance of accuracy and latency, achieving utterance F1 of 0.93 with 7.8% false alarms at 400ms median latency. Adding text increases premature detections without improving performance. Feature space analysis confirms that prosodic features have the strongest class separability, while text rep
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית