יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

GEPARD - דגם טקסט-קול לדיאלוג באורך זמן

GEPARD - Generative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue
GEPARD - דגם טקסט-קול לדיאלוג באורך זמן. GEPARD מפיק דיבור אוטורגרסיבי עם רקע LLM - המסרים והאמפתיזציות הקוליות הוכשרו יחד בדגם רק-מקדם-אחד.
תקציר מקורי באנגליתarXiv:2609.04222v1 Announce Type: cross Abstract: We present GEPARD (Generative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue), a streaming text-to-speech model for real-time spoken dialogue. GEPARD generates speech autoregressively with an LLM backbone - text and audio embeddings are trained together in a single decoder-only model - and decodes it to a waveform with an FSQ-based neural codec, streaming audio chunk-by-chunk as text arrives. Our central goal is a TTS architecture served by a standard LLM engine (vLLM) without modifying its compute kernels. This defines the overarching design principle: the backbone is a standard full-attention transformer, while all non-trivial auxiliary mechanisms - zero-shot voice cloning, text augmentation, and classifier-free
קרא במקור המקורי