יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

הנעה רגשית במודל שיחה דו-כיוונית

Steering Follows Geometry, Not Labels: Emotion Directions in a Full-Duplex Speech Model
חוקרים בדו-כיוונית מודל שיחה Moshi, המאפשר שליטה על רגשות ואופן מסירה בזמן אמת. המחקר מראה כי ניתן לקדוד רגשות באופן ליניארי מזרם השאריות של Moshi, אך הנעה פעילה חלקית בלבד. הממצאים מצביעים על כך שרגשות שונים נעים לכיוונים שונים.
תקציר מקורי באנגליתarXiv:2610.08887v1 Announce Type: new Abstract: Full-duplex voice agents need to modulate emotion and delivery during real-time conversations, when de-escalating a complaint, carrying urgency in dispatch, softening a clinical result. Emotion and delivery control is well studied for TTS and turn based models through prompt-conditioned synthesis, reference-conditioned synthesis and activation steering; PersonaPlex controls identity in a duplex model but not affect. We study emotion steering in Moshi, a fully open sourced full-duplex speech language model, across four emotions, using mean-difference activation steering, which costs only a few vector additions per frame and no retraining. We show that emotion is linearly decodable from Moshi's residual stream, but activation steering is only p
קרא במקור המקורי