Fish Voice Cloner
Clone a voice from a short recording and speak with it in 13 languages.
Fish Voice Clone
Clone a voice from a short recording and speak with it in 13 languages, right in your browser on mitte.ai.
Fish Voice Clone turns a short audio clip into a reusable voice. You describe the voice, upload a sample, optionally add a thumbnail, and a few seconds later the voice sits in your presets, ready to narrate anything you type. It runs on Fish Audio S2, and on mitte.ai there’s nothing to install and no code to write.
What is voice cloning?
Voice cloning builds a synthetic copy of a real voice from a recording. The model listens to a short sample and picks up the things that make a voice recognisable: timbre, cadence, the small hesitations and inflections between words. Then it can read any text in that voice.
Older systems needed studio sessions and hours of training. Fish Audio S2 works from about 10 seconds of audio and returns a usable voice in seconds. Emotion carries through the clone, so a dry read stays dry and a warm read stays warm.
The clone is also cross-lingual. Record a sample in English and the voice can speak Spanish, Japanese, Arabic and 10 more languages with the same character, no re-recording needed.
Key features
- A short sample is enough: 10 to 30 seconds of clear speech builds a working voice.
- Ready in seconds: no training queue. The voice is usable almost as soon as it’s created.
- Keeps the character: timbre, pacing and emotional delivery survive the clone, so the result sounds like the person, not a generic narrator.
- 13 languages from one clip: clone once, speak everywhere. The voice keeps its identity across languages.
- Handles rough recordings: background noise gets cleaned up automatically, so a phone recording in a quiet room works fine.
- Saved as a preset: every voice you clone lives under My presets, private to you, ready to reuse in any project.
How to clone a voice on mitte.ai
- Open the Voice Cloner on mitte.ai.
- Describe the voice. Give it a name and note the language, accent or tone, e.g. “Amélie, French, warm and unhurried.”
- Upload a voice sample. 10 to 30 seconds of clear speech. An audio file works, and so does an mp4 video; the voice is pulled from its audio.
- Add a thumbnail if you want a cover image for the voice. This step is optional.
- Run it. The voice appears under My presets a few seconds later.
- Use it. Open the preset, paste your script, pick a language and generate speech.
How to control emotions?
Once the voice is cloned, you direct its delivery from the text itself. When you generate speech with the preset, place cues in square brackets where you want the performance to change:
[happy] Welcome back to the show!
[whispering] Don't tell anyone yet.
[sad][sighing] I wish things were different. Sigh.
A tag at the start of a sentence sets its emotion; you can layer tags for combined effects, add human sounds like [laughing] and [gasping], or write free-form descriptions like [slightly nervous]. There are 64+ expressions in total. See the full tag reference on the Fish Text-to-speech page.
Getting a good sample
The clone is only as good as the recording you feed it. A few things help:
- One speaker, no overlap. The model should hear a single voice.
- No music or reverb. A quiet room beats a bathroom or a café.
- Natural delivery. Speak the way you’d want the clone to speak. If you want warmth, record warmth.
- 10 to 30 seconds. Enough to capture the voice; varied, natural speech beats a monotone read.
Any common audio format works, and you can drop in an mp4 video instead; the sample is taken from its audio track.
What creatives use it for
- Narration in your own voice. Record once, then generate voiceover for cuts, edits and revisions without booking yourself into a booth every time the script changes.
- Character voices. Build distinct voices for animation, games and story work, and keep them consistent across episodes.
- Localization. Ship the same campaign in Spanish, Japanese or German with the original speaker’s voice intact.
- Scratch VO. Drop temp narration into storyboards and animatics that sounds close enough to sell the edit.
- A consistent brand voice. One voice across product videos, ads and tutorials, without depending on one person’s calendar.
FAQ
How much audio do I need to clone a voice? 10 to 30 seconds of clear, single-speaker speech. There’s no need for long studio sessions.
How long does it take? Seconds. There’s no training queue; the voice is usable almost immediately after it’s created.
Can the clone speak languages the original speaker doesn’t? Yes. Fish Audio S2 speaks 13 languages from a single sample, keeping the voice’s character across all of them.
Where does my cloned voice live? Under My presets on mitte.ai, private to your account. Reuse it in any project by opening the preset and typing new text.
Do I need to write code or set up an API? No. Everything happens in the browser: describe the voice, upload a sample, run, generate.
Whose voice can I clone? Your own, or a voice you have clear permission to use. Cloning someone without their consent can cross legal and ethical lines, so get it in writing for client and commercial work.
What model powers it? Fish Audio S2, a speech model built for fast cloning and natural, emotionally faithful output.
Ready to hear yourself? Clone your voice on mitte.ai.