Fish Text-to-speech
Turn up to 50,000 characters of text into lifelike speech, with 64+ emotion and delivery tags.
Fish Text-to-speech
Turn up to 50,000 characters of text into lifelike speech, with 64+ emotion and delivery tags, in your browser on mitte.ai.
Fish Text-to-speech reads your script aloud in a natural voice. Paste text, pick a model, generate. The result lands in your library as an audio asset, ready to drop into an edit. It runs on Fish Audio’s S2.1 Pro model, and you direct the performance with simple inline tags: [happy], [whispering], [sighing].
What is Fish Text-to-speech?
Text to speech converts written words into spoken audio. Fish Audio’s S2 models go further than a flat read: they carry emotion, pacing and emphasis, and you control all of it from the text itself. Write [excited] We won! and the voice sounds excited. Write [sad][whispering] I miss you and it drops to a sorrowful whisper.
The voice itself is up to you. Use the default voice, or speak in a voice you’ve cloned with the Voice Cloner; cloned voices live under My presets and keep their character across all 13 languages.
Key features
- Long-form input: up to 50,000 characters per run, enough for a chapter, a full explainer script or an episode of narration.
- Three models: S2.1 Pro is the default and the best quality; S2 Pro and S1 are there when you need to match older work.
- 64+ emotion and delivery tags: happy, sarcastic, whispering, sighing, audience laughter and more, all controlled inline from the text.
- Your own voice: pair it with a cloned voice preset and narrate in your voice without recording a word.
- Speed control: from 0.5x to 2x, set in the additional settings.
- Format options: MP3 at 64, 128 or 192 kbps, WAV, or Opus.
How to generate speech on mitte.ai
- Open Fish Text-to-speech on mitte.ai.
- Paste your text. Up to 50,000 characters. Add emotion tags where you want the delivery to change (see below).
- Pick a model. S2.1 Pro is selected by default and is the right choice for most work.
- Adjust the settings if you need to: speed, output format, bitrate, or a specific voice ID.
- Generate. The audio appears in your library, ready to use.
To speak in your own voice, clone it first and run it from the preset under My presets.
Directing the performance: emotion tags
The S2 models read cues you place in square brackets. A tag at the start of a sentence sets the emotion for that sentence; tone and sound tags can go anywhere in the text.
[happy] What a beautiful day!
[sad] I'm sorry to hear that.
[excited] This is amazing news!
The brackets accept free-form natural language too, so you can write descriptions instead of picking from a fixed list:
What a [warm and happy] wonderful day!
[slightly sad] I'm a bit disappointed.
[very excited] This is absolutely amazing!
[extremely angry] This is unacceptable!
Basic emotions
| Tag | Description | Example context |
|---|---|---|
[happy] |
Cheerful, upbeat tone | Good news, greetings |
[sad] |
Melancholic, downcast | Sympathy, bad news |
[angry] |
Frustrated, aggressive | Complaints, warnings |
[excited] |
Energetic, enthusiastic | Announcements, celebrations |
[calm] |
Peaceful, relaxed | Instructions, meditation |
[nervous] |
Anxious, uncertain | Disclaimers, apologies |
[confident] |
Assertive, self-assured | Presentations, sales |
[surprised] |
Shocked, amazed | Reactions, discoveries |
[satisfied] |
Content, pleased | Confirmations, reviews |
[delighted] |
Very pleased, joyful | Celebrations, compliments |
[scared] |
Frightened, fearful | Warnings, horror stories |
[worried] |
Concerned, troubled | Concerns, questions |
[upset] |
Disturbed, distressed | Complaints, problems |
[frustrated] |
Annoyed, exasperated | Technical issues, delays |
[depressed] |
Very sad, hopeless | Serious topics |
[empathetic] |
Understanding, caring | Support, counseling |
[embarrassed] |
Ashamed, awkward | Apologies, mistakes |
[disgusted] |
Repelled, revolted | Negative reviews |
[moved] |
Emotionally touched | Heartfelt moments |
[proud] |
Accomplished, satisfied | Achievements, praise |
[relaxed] |
At ease, casual | Casual conversation |
[grateful] |
Thankful, appreciative | Thanks, appreciation |
[curious] |
Inquisitive, interested | Questions, exploration |
[sarcastic] |
Ironic, mocking | Humor, criticism |
Advanced emotions
| Tag | Description | Example context |
|---|---|---|
[disdainful] |
Contemptuous, scornful | Criticism, rejection |
[unhappy] |
Discontent, dissatisfied | Complaints, feedback |
[anxious] |
Very worried, uneasy | Urgent matters |
[hysterical] |
Uncontrollably emotional | Extreme reactions |
[indifferent] |
Uncaring, neutral | Neutral responses |
[uncertain] |
Doubtful, unsure | Speculation, questions |
[doubtful] |
Skeptical, questioning | Disbelief, questioning |
[confused] |
Puzzled, perplexed | Clarification requests |
[disappointed] |
Let down, dissatisfied | Unmet expectations |
[regretful] |
Sorry, remorseful | Apologies, mistakes |
[guilty] |
Culpable, responsible | Confessions, apologies |
[ashamed] |
Deeply embarrassed | Serious mistakes |
[jealous] |
Envious, resentful | Comparisons |
[envious] |
Wanting what others have | Admiration with desire |
[hopeful] |
Optimistic about future | Future plans |
[optimistic] |
Positive outlook | Encouragement |
[pessimistic] |
Negative outlook | Warnings, doubts |
[nostalgic] |
Longing for the past | Memories, stories |
[lonely] |
Isolated, alone | Emotional content |
[bored] |
Uninterested, weary | Disinterest |
[contemptuous] |
Showing contempt | Strong criticism |
[sympathetic] |
Showing sympathy | Condolences |
[compassionate] |
Showing deep care | Support, help |
[determined] |
Resolved, decided | Goals, commitments |
[resigned] |
Accepting defeat | Giving up, acceptance |
Tone markers
These shape volume, intensity and emphasis. Place [emphasis] right before the word or phrase you want stressed:
This is [emphasis] really important.
| Tag | Description | When to use |
|---|---|---|
[in a hurry tone] |
Rushed, urgent | Time-sensitive information |
[shouting] |
Loud, calling out | Getting attention |
[screaming] |
Very loud, panicked | Emergencies, fear |
[whispering] |
Very soft, secretive | Secrets, quiet scenes |
[soft tone] |
Gentle, quiet | Comfort, lullabies |
[emphasis] |
Stress a word or phrase | Highlighting key words |
Human sounds
Add natural, non-verbal sounds. Give the sound something to work with by adding matching text after the tag (“Ha, ha” after laughing):
| Tag | Description | Suggested text |
|---|---|---|
[laughing] |
Full laughter | Ha, ha, ha |
[chuckling] |
Light laugh | Heh, heh |
[sobbing] |
Crying heavily | Optional text |
[crying loudly] |
Intense crying | Optional text |
[sighing] |
Exhale of relief or frustration | sigh |
[groaning] |
Sound of frustration | ugh |
[panting] |
Out of breath | huff, puff |
[gasping] |
Sharp intake of breath | gasp |
[yawning] |
Tired sound | yawn |
[snoring] |
Sleep sound | zzz |
[clear throat] |
Throat-clearing sound | ahem |
You can also write natural expressions like “Ha, ha, ha” for laughter without any tag.
Special effects and pauses
| Tag | Description |
|---|---|
[audience laughing] |
Crowd laughing sound |
[background laughter] |
Ambient laughter |
[crowd laughing] |
Large group laughing |
[break] |
Brief pause in speech |
[long-break] |
Extended pause in speech |
Combining tags
Layer up to 3 tags per sentence for complex delivery:
[sad][whispering] I miss you so much.
[angry][shouting] Get out of here now!
[excited][laughing] We won! Ha ha!
And walk a passage through an emotional arc, one cue per sentence:
[happy] I got the promotion!
[uncertain] But... it means relocating.
[sad] I'll miss everyone here.
[hopeful] Though it's a great opportunity.
[determined] I'm going to make it work!
Rough intensity ladder, when you want to dial an emotion up or down:
| Base emotion | Mild | Moderate | Intense |
|---|---|---|---|
| Happy | satisfied | happy | delighted |
| Sad | disappointed | sad | depressed |
| Angry | frustrated | angry | furious |
| Scared | nervous | scared | terrified |
| Excited | interested | excited | ecstatic |
Combinations that reliably work:
| Scenario | Tags | Example |
|---|---|---|
| Whispered secret |
[mysterious][whispering] |
“I have something to tell you…” |
| Angry shout |
[angry][shouting] |
“Stop right there!” |
| Sad sigh |
[sad][sighing] |
“I wish things were different. Sigh.” |
| Excited laugh |
[excited][laughing] |
“We did it! Ha ha!” |
| Nervous question |
[nervous][uncertain] |
“Are you sure about this?” |
Getting tags to behave
- Put sentence-level emotion cues at the start of the sentence they control; keep them close to the words they affect.
- Use one primary emotion per sentence, and space emotional changes out for realism.
- Match emotions to context logically, and don’t mix conflicting emotions in one sentence.
- Keep free-form descriptions short; a bracket that runs half a line hurts readability and the read.
- Don’t forget the brackets, and don’t pile tags into short text.
- Tags work in all 13 languages: English, Chinese, Japanese, German, French, Spanish, Korean, Arabic, Russian, Dutch, Italian, Polish and Portuguese.
-
On the legacy S1 model, wrap the same tags in parentheses instead of brackets, e.g.
(happy)or(sad)(whispering). S1 only understands its fixed tag set; free-form descriptions and intensity modifiers are an S2 feature.
Exact pronunciation and speech rhythm
For names, brand names and technical terms the model reads wrong, you can spell out the exact pronunciation with phoneme tags. Wrap the pronunciation in <|phoneme_start|> and <|phoneme_end|>; how you write it depends on the language:
-
English: replace one word with CMU Arpabet.
I am an <|phoneme_start|>EH1 N JH AH0 N IH1 R<|phoneme_end|>. -
Chinese: replace one character or syllable with tone-number pinyin, one tag per syllable.
我是一个<|phoneme_start|>gong1<|phoneme_end|><|phoneme_start|>cheng2<|phoneme_end|><|phoneme_start|>shi1<|phoneme_end|>。 -
Japanese: replace a short word or phrase with OpenJTalk-style romaji plus pitch digits (0 low, 1 high).
<|phoneme_start|>ha0shi1ga0<|phoneme_end|>見えます。The pitch digits disambiguate homographs like 端, 箸 and 橋, which share the same sounds.
Keep punctuation outside the tags, and tag short spans rather than whole paragraphs.
For rhythm, write natural pause words like “um” and “uh” straight into the text, or use parenthesis effects: (break) for a short pause, (long-break) for a longer one, plus (breath), (laugh), (cough), (lip-smacking) and (sigh). The last four are experimental and may need repeating to land.
I am, um, an (break) <|phoneme_start|>EH1 N JH AH0 N IH1 R<|phoneme_end|>.
What creatives use it for
- Narration. Audiobooks, documentaries and explainers, with emotion tags doing the vocal direction a session with a voice actor would.
- Character dialogue. Distinct deliveries for animation and game scripts; layer emotions per line and keep takes consistent across revisions.
- Video voiceover. Scratch or final VO for edits, ads and social cuts, regenerated in seconds every time the script changes.
- Localization. The same script and the same voice across 13 languages.
- Your voice, at scale. Clone your voice once and narrate everything in it without booking yourself into a booth.
FAQ
How much text can I convert at once? Up to 50,000 characters per run, enough for very long scripts and full chapters.
Which model should I pick? S2.1 Pro, the default. It’s the latest and highest quality. S2 Pro and S1 are previous generations, useful when you need to match audio you generated with them.
Can it speak in my voice? Yes. Clone your voice with the Voice Cloner first; the cloned voice appears under My presets and you generate speech from there. The additional settings also accept a Fish Audio voice ID directly.
How do I control emotion?
With inline tags in square brackets: [happy], [whispering], [sarcastic] and 60+ more, including free-form descriptions like [warm and happy]. See the full reference above.
Which languages does it support? 13: English, Chinese, Japanese, German, French, Spanish, Korean, Arabic, Russian, Dutch, Italian, Polish and Portuguese. Emotion tags work in all of them.
What audio formats can it output? MP3 (at 64, 128 or 192 kbps), WAV, or Opus. MP3 at 128 kbps is the default.
Can I change the speaking speed? Yes, from 0.5x to 2x in the additional settings.
Got a script waiting? Give it a voice on mitte.ai.