כתבה
arXiv cs.AI ·
NemotronLabs VoiceChat: מודל דיבור-לדיבור פתוח עם יכולת קריאה לכלים
NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
מודל דיבור-לדיבור פתוח שמאפשר קריאה לכלים, תרגום, ודיבור. המודל משלב קודקס דיבור-לדיבור, קודקס טקסט, וקודקס קריאה לכלים. המודל עובד בזמן אמת ומסוגל לשמור על תקשורת טבעית.
תקציר מקורי באנגליתarXiv:2609.21967v2 Announce Type: replace-cross Abstract: We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder and decoder-only language model with parallel specialized output streams for agent text and structured function calls, an auxiliary RNN-T branch for incremental user transcription, and a streaming TTS decoder. This design enables the model to listen, transcribe, reason, invoke tools, and speak within a unified streaming architecture while preserving the temporal behavior required for natural conversation. On Full-Duplex-Bench 1.0, NemotronLabs VoiceChat achieves the lowest pause-handling takeover rates among evaluated open-weight systems, 100\% takeover foll
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית