Voice Cloner

Clone a voice from a short recording and speak with it in 13 languages — emotion carries through the clone. Runs on Fish Audio's S2 model.

AI Voice Cloner

Clone a voice from a short recording and speak with it in 13 languages.

Upload a short audio sample and this tool builds a reusable voice from it — no studio session, no lengthy training. The clone captures timbre, cadence, and delivery, and it’s ready to narrate anything you type in seconds. It runs on Fish Audio’s S2 model.

How to clone a voice

  1. Upload a clean, single-speaker recording — 10 to 30 seconds is enough.
  2. Optionally add a name and thumbnail for the voice.
  3. Generate. The voice lands in your presets, ready to use with Text to Speech.

Emotion control

Once cloned, the voice speaks with the same inline emotion tags as Text to Speech — the clone doesn’t just copy a sound, it copies how that sound carries emotion:

See the full tag reference and advanced usage on the node page for the complete list of 64+ tags.

Getting a good sample

What creatives use it for

FAQ

How does voice cloning work?

The model listens to your sample and learns the traits that make a voice recognizable — timbre, cadence, small inflections — then uses that profile to read any new text you give it.

How do I clone a voice with AI?

Upload 10 to 30 seconds of a clean, single-speaker recording. There’s no editing or training wait — the voice is ready to use within seconds.

Is AI voice cloning legal?

Only clone voices you have the right to use — your own voice, or a voice you have explicit permission for. Using someone else’s voice without consent can violate their rights and, depending on the use, the law.

What’s the difference between this and the Clone Voice preset?

Clone Voice uses a different underlying model. This tool runs on Fish Audio S2, which carries emotion through the clone and pairs directly with Text to Speech‘s full emotion-tag system for expressive delivery.

All apps

See pricing and start creating on Mitte