Voice Cloner
Clone a voice from a short recording and speak with it in 13 languages — emotion carries through the clone. Runs on Fish Audio's S2 model.
AI Voice Cloner
Clone a voice from a short recording and speak with it in 13 languages.
Upload a short audio sample and this tool builds a reusable voice from it — no studio session, no lengthy training. The clone captures timbre, cadence, and delivery, and it’s ready to narrate anything you type in seconds. It runs on Fish Audio’s S2 model.
How to clone a voice
- Upload a clean, single-speaker recording — 10 to 30 seconds is enough.
- Optionally add a name and thumbnail for the voice.
- Generate. The voice lands in your presets, ready to use with Text to Speech.
Emotion control
Once cloned, the voice speaks with the same inline emotion tags as Text to Speech — the clone doesn’t just copy a sound, it copies how that sound carries emotion:
-
Wrap an emotion in brackets before the words it should affect:
[happy],[sad],[whispering]. -
Stack tags for nuance, like
[angry][shouting], and adjust intensity with words like “slightly” or “very”.
See the full tag reference and advanced usage on the node page for the complete list of 64+ tags.
Getting a good sample
- One speaker, no overlap. The model should hear a single voice.
- No music or reverb. A quiet room beats a noisy one.
- Natural delivery. Record the way you want the clone to sound.
- 10 to 30 seconds. Enough to capture the voice — you can also drop in a video and the audio track is used.
What creatives use it for
- Narration in your own voice. Re-record a script without booking a booth every time it changes.
- Character voices. Build distinct voices for animation, games, and story work.
- Localization. Ship the same campaign in Spanish, Japanese, or German with the original speaker’s voice intact.
- A consistent brand voice. One voice across product videos, ads, and tutorials.
FAQ
How does voice cloning work?
The model listens to your sample and learns the traits that make a voice recognizable — timbre, cadence, small inflections — then uses that profile to read any new text you give it.
How do I clone a voice with AI?
Upload 10 to 30 seconds of a clean, single-speaker recording. There’s no editing or training wait — the voice is ready to use within seconds.
Is AI voice cloning legal?
Only clone voices you have the right to use — your own voice, or a voice you have explicit permission for. Using someone else’s voice without consent can violate their rights and, depending on the use, the law.
What’s the difference between this and the Clone Voice preset?
Clone Voice uses a different underlying model. This tool runs on Fish Audio S2, which carries emotion through the clone and pairs directly with Text to Speech‘s full emotion-tag system for expressive delivery.