כתבה
arXiv cs.AI ·
Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models
תקציר מקורי באנגליתarXiv:2609.09263v1 Announce Type: cross Abstract: Speech-to-speech (S2S) models now run inside dubbing, translation, and voice agents. Unlike text models, they hear the speaker's voice, which carries the speaker's gender. A faithful system should treat a speaker as who they sound like, not as whoever usually says what they said. Testing this is harder than it looks, since most S2S models answer in a single, fixed output voice, hard-coded so it cannot drift toward a stereotype. Checking the output voice comes back clean even when the model is biased. We therefore ask two questions. When a model re-speaks the input, does the stereotype in the words shift the perceived gender of the output voice (voice rendering)? And when the model states the speaker's gender, does it follow the voice or the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית