Voice Conversion: How Speech-to-Speech Voice Changing Works
Voice conversion (also called speech-to-speech) takes an existing recording and re-renders it in a different target voice while keeping the original timing, emotion, and delivery. You use it when the performance already exists and you only want to change who it sounds like — not what is said or how it is paced. If you instead need to start from written text, you want text-to-speech; if you need a reusable custom voice you can apply again and again, you want voice cloning. Uberduck lists voice conversion as an API-accessible capability alongside text-to-speech, voice cloning, and speech-to-speech, which makes it a concrete example of where this fits in a real toolset.
Voice conversion vs. text-to-speech vs. voice cloning
These three get conflated constantly, but they start from different inputs and solve different problems.
| Input | What it produces | Best for | |
|---|---|---|---|
| Voice conversion (speech-to-speech) | An existing audio recording | The same performance in a different voice | Changing the speaker on a take you already like |
| Text-to-speech | Written text | New speech generated from scratch | Turning scripts, articles, or UI copy into audio |
| Voice cloning | Voice samples | A reusable custom voice model | Building a voice you can apply to many future projects |
The practical test: if you already have audio and want to keep its rhythm and emotion, that is voice conversion. If you have nothing but words on a page, that is text-to-speech. If your goal is to own a voice you can reuse, that is cloning — and cloning often feeds into the other two as the target voice.
The typical voice conversion workflow
The steps are consistent across tools, even when the interface differs:
- Provide source audio. Upload or record the performance you want to convert. Clean, single-speaker audio with minimal background noise converts far better than a noisy phone recording.
- Choose or create a target voice. This can be a built-in voice or one you created through cloning. The target determines what the output sounds like.
- Run the conversion. The system maps the source performance onto the target voice, preserving timing and delivery.
- Review and export. Listen for artifacts, check that emotion and pacing survived, then export.
Uberduck exposes this through its API, so the same flow can be scripted rather than done by hand — useful if you are converting many clips or wiring it into a larger pipeline.
What you need for a usable result
- Source audio quality. Clear, single-speaker input is the single biggest factor. Overlapping speakers, heavy reverb, or clipping will carry through.
- Language support. Voice conversion inherits the language constraints of the tool. Uberduck supports 70+ languages for its speech features, so check that your source language is covered before committing.
- Rights and consent. Converting a recording into someone else's voice — especially a real person's — raises consent and licensing questions. Only use voices you have the right to use, and be careful with public figures.
Common failure cases and how to troubleshoot
- Robotic or metallic artifacts. Usually a sign of low-quality source audio or an aggressive conversion. Try a cleaner source clip first.
- Mismatched accent or tone. The target voice may not match the source language or delivery style. Test a few target voices rather than forcing one.
- Unclear source audio. If the original is muffled or noisy, the conversion amplifies the problem. Re-record or clean the audio before converting.
- Emotion that feels flat. Some conversions trade expressiveness for stability. Compare a short test clip before converting a long piece.
Where voice conversion fits
Reach for voice conversion when the performance is already right and only the speaker needs to change: dubbing a take into a different character voice, adapting a scratch recording into a final voice, or restyling existing audio without re-recording. Uberduck is one tool that offers it as part of a broader AI vocals platform, which is convenient if you also need text-to-speech, singing, or cloning in the same project. If you only ever start from text, you can skip voice conversion entirely and go straight to text-to-speech.