וידאו
YT AI Engineer ·
Speech-to-Speech Model Research at Google DeepMind — Valeria Wu Fon & Tom Ouyang, Google DeepMind
▶ צפה כאן — בלי לצאת מהאתר
תקציר מקורי באנגליתAsked in Spanish about a mid century sofa, the model answers in Spanish but leaves the phrase mid century in English, because that is how the term is actually used by Spanish speakers. Nobody wrote a rule for that. It falls out of a model trained on audio, video, and text together rather than stitched from separate parts. Valeria Wu Fon leads product for speech to speech in Gemini and Tom Ouyang engineers it, and they open with how far the field moved to get here. Before roughly 2018, turning speech into text meant a chain of hand built components: feature extraction, acoustic modeling, pronunciation modeling, language modeling, and a rescoring pass. End to end models collapsed that chain but still only did the one job. Ask for the speaker's tone, or emotion, or pace, and you were back to
קרא במקור המקורי
youtube.com
פתח כתבה מקורית