יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

FD-VAD: זיהוי סמנטי של סיום דיבור בשיחה דו-כיוונית

FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech
פיתוח זיהוי סמנטי של סיום דיבור בשיחה דו-כיוונית, המאפשר זיהוי סיום דיבור על ידי שימוש במודלי שפה ובעזרת רשתות חישוביות. המחקר עוסק בפיתוח שיטה חדשה לזיהוי סיום דיבור, המבוססת על שימוש במודלי שפה ובעזרת רשתות חישוביות. השיטה נועדה לשפר את יכולת הזיהוי של סיום דיבור בשיחה דו-כיוונית, ולהפחית את הזמן הדרוש לזיהוי.
תקציר מקורי באנגליתarXiv:2609.35791v1 Announce Type: new Abstract: Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dependence and additional processing stages. We formulate semantic endpoint detection as a causal audio-language reasoning task and introduce FD-VAD, an ASR-free streaming endpointer that maps bounded causal audio windows directly to Continue/Stop decisions. FD-VAD combines a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model, using a last-chunk training objective for streaming inference. We fu
קרא במקור המקורי