Fish Text-to-speech

Turn up to 50,000 characters of text into lifelike speech, with 64+ emotion and delivery tags.

Fish Text-to-speech

Turn up to 50,000 characters of text into lifelike speech, with 64+ emotion and delivery tags, in your browser on mitte.ai.

Fish Text-to-speech reads your script aloud in a natural voice. Paste text, pick a model, generate. The result lands in your library as an audio asset, ready to drop into an edit. It runs on Fish Audio’s S2.1 Pro model, and you direct the performance with simple inline tags: [happy], [whispering], [sighing].

What is Fish Text-to-speech?

Text to speech converts written words into spoken audio. Fish Audio’s S2 models go further than a flat read: they carry emotion, pacing and emphasis, and you control all of it from the text itself. Write [excited] We won! and the voice sounds excited. Write [sad][whispering] I miss you and it drops to a sorrowful whisper.

The voice itself is up to you. Use the default voice, or speak in a voice you’ve cloned with the Voice Cloner; cloned voices live under My presets and keep their character across all 13 languages.

Key features

  • Long-form input: up to 50,000 characters per run, enough for a chapter, a full explainer script or an episode of narration.
  • Three models: S2.1 Pro is the default and the best quality; S2 Pro and S1 are there when you need to match older work.
  • 64+ emotion and delivery tags: happy, sarcastic, whispering, sighing, audience laughter and more, all controlled inline from the text.
  • Your own voice: pair it with a cloned voice preset and narrate in your voice without recording a word.
  • Speed control: from 0.5x to 2x, set in the additional settings.
  • Format options: MP3 at 64, 128 or 192 kbps, WAV, or Opus.

How to generate speech on mitte.ai

  1. Open Fish Text-to-speech on mitte.ai.
  2. Paste your text. Up to 50,000 characters. Add emotion tags where you want the delivery to change (see below).
  3. Pick a model. S2.1 Pro is selected by default and is the right choice for most work.
  4. Adjust the settings if you need to: speed, output format, bitrate, or a specific voice ID.
  5. Generate. The audio appears in your library, ready to use.

To speak in your own voice, clone it first and run it from the preset under My presets.

Directing the performance: emotion tags

The S2 models read cues you place in square brackets. A tag at the start of a sentence sets the emotion for that sentence; tone and sound tags can go anywhere in the text.

[happy] What a beautiful day!
[sad] I'm sorry to hear that.
[excited] This is amazing news!

The brackets accept free-form natural language too, so you can write descriptions instead of picking from a fixed list:

What a [warm and happy] wonderful day!
[slightly sad] I'm a bit disappointed.
[very excited] This is absolutely amazing!
[extremely angry] This is unacceptable!

Basic emotions

Tag Description Example context
[happy] Cheerful, upbeat tone Good news, greetings
[sad] Melancholic, downcast Sympathy, bad news
[angry] Frustrated, aggressive Complaints, warnings
[excited] Energetic, enthusiastic Announcements, celebrations
[calm] Peaceful, relaxed Instructions, meditation
[nervous] Anxious, uncertain Disclaimers, apologies
[confident] Assertive, self-assured Presentations, sales
[surprised] Shocked, amazed Reactions, discoveries
[satisfied] Content, pleased Confirmations, reviews
[delighted] Very pleased, joyful Celebrations, compliments
[scared] Frightened, fearful Warnings, horror stories
[worried] Concerned, troubled Concerns, questions
[upset] Disturbed, distressed Complaints, problems
[frustrated] Annoyed, exasperated Technical issues, delays
[depressed] Very sad, hopeless Serious topics
[empathetic] Understanding, caring Support, counseling
[embarrassed] Ashamed, awkward Apologies, mistakes
[disgusted] Repelled, revolted Negative reviews
[moved] Emotionally touched Heartfelt moments
[proud] Accomplished, satisfied Achievements, praise
[relaxed] At ease, casual Casual conversation
[grateful] Thankful, appreciative Thanks, appreciation
[curious] Inquisitive, interested Questions, exploration
[sarcastic] Ironic, mocking Humor, criticism

Advanced emotions

Tag Description Example context
[disdainful] Contemptuous, scornful Criticism, rejection
[unhappy] Discontent, dissatisfied Complaints, feedback
[anxious] Very worried, uneasy Urgent matters
[hysterical] Uncontrollably emotional Extreme reactions
[indifferent] Uncaring, neutral Neutral responses
[uncertain] Doubtful, unsure Speculation, questions
[doubtful] Skeptical, questioning Disbelief, questioning
[confused] Puzzled, perplexed Clarification requests
[disappointed] Let down, dissatisfied Unmet expectations
[regretful] Sorry, remorseful Apologies, mistakes
[guilty] Culpable, responsible Confessions, apologies
[ashamed] Deeply embarrassed Serious mistakes
[jealous] Envious, resentful Comparisons
[envious] Wanting what others have Admiration with desire
[hopeful] Optimistic about future Future plans
[optimistic] Positive outlook Encouragement
[pessimistic] Negative outlook Warnings, doubts
[nostalgic] Longing for the past Memories, stories
[lonely] Isolated, alone Emotional content
[bored] Uninterested, weary Disinterest
[contemptuous] Showing contempt Strong criticism
[sympathetic] Showing sympathy Condolences
[compassionate] Showing deep care Support, help
[determined] Resolved, decided Goals, commitments
[resigned] Accepting defeat Giving up, acceptance

Tone markers

These shape volume, intensity and emphasis. Place [emphasis] right before the word or phrase you want stressed:

This is [emphasis] really important.
Tag Description When to use
[in a hurry tone] Rushed, urgent Time-sensitive information
[shouting] Loud, calling out Getting attention
[screaming] Very loud, panicked Emergencies, fear
[whispering] Very soft, secretive Secrets, quiet scenes
[soft tone] Gentle, quiet Comfort, lullabies
[emphasis] Stress a word or phrase Highlighting key words

Human sounds

Add natural, non-verbal sounds. Give the sound something to work with by adding matching text after the tag (“Ha, ha” after laughing):

Tag Description Suggested text
[laughing] Full laughter Ha, ha, ha
[chuckling] Light laugh Heh, heh
[sobbing] Crying heavily Optional text
[crying loudly] Intense crying Optional text
[sighing] Exhale of relief or frustration sigh
[groaning] Sound of frustration ugh
[panting] Out of breath huff, puff
[gasping] Sharp intake of breath gasp
[yawning] Tired sound yawn
[snoring] Sleep sound zzz
[clear throat] Throat-clearing sound ahem

You can also write natural expressions like “Ha, ha, ha” for laughter without any tag.

Special effects and pauses

Tag Description
[audience laughing] Crowd laughing sound
[background laughter] Ambient laughter
[crowd laughing] Large group laughing
[break] Brief pause in speech
[long-break] Extended pause in speech

Combining tags

Layer up to 3 tags per sentence for complex delivery:

[sad][whispering] I miss you so much.
[angry][shouting] Get out of here now!
[excited][laughing] We won! Ha ha!

And walk a passage through an emotional arc, one cue per sentence:

[happy] I got the promotion!
[uncertain] But... it means relocating.
[sad] I'll miss everyone here.
[hopeful] Though it's a great opportunity.
[determined] I'm going to make it work!

Rough intensity ladder, when you want to dial an emotion up or down:

Base emotion Mild Moderate Intense
Happy satisfied happy delighted
Sad disappointed sad depressed
Angry frustrated angry furious
Scared nervous scared terrified
Excited interested excited ecstatic

Combinations that reliably work:

Scenario Tags Example
Whispered secret [mysterious][whispering] “I have something to tell you…”
Angry shout [angry][shouting] “Stop right there!”
Sad sigh [sad][sighing] “I wish things were different. Sigh.”
Excited laugh [excited][laughing] “We did it! Ha ha!”
Nervous question [nervous][uncertain] “Are you sure about this?”

Getting tags to behave

  • Put sentence-level emotion cues at the start of the sentence they control; keep them close to the words they affect.
  • Use one primary emotion per sentence, and space emotional changes out for realism.
  • Match emotions to context logically, and don’t mix conflicting emotions in one sentence.
  • Keep free-form descriptions short; a bracket that runs half a line hurts readability and the read.
  • Don’t forget the brackets, and don’t pile tags into short text.
  • Tags work in all 13 languages: English, Chinese, Japanese, German, French, Spanish, Korean, Arabic, Russian, Dutch, Italian, Polish and Portuguese.
  • On the legacy S1 model, wrap the same tags in parentheses instead of brackets, e.g. (happy) or (sad)(whispering). S1 only understands its fixed tag set; free-form descriptions and intensity modifiers are an S2 feature.

Exact pronunciation and speech rhythm

For names, brand names and technical terms the model reads wrong, you can spell out the exact pronunciation with phoneme tags. Wrap the pronunciation in <|phoneme_start|> and <|phoneme_end|>; how you write it depends on the language:

  • English: replace one word with CMU Arpabet. I am an <|phoneme_start|>EH1 N JH AH0 N IH1 R<|phoneme_end|>.
  • Chinese: replace one character or syllable with tone-number pinyin, one tag per syllable. 我是一个<|phoneme_start|>gong1<|phoneme_end|><|phoneme_start|>cheng2<|phoneme_end|><|phoneme_start|>shi1<|phoneme_end|>。
  • Japanese: replace a short word or phrase with OpenJTalk-style romaji plus pitch digits (0 low, 1 high). <|phoneme_start|>ha0shi1ga0<|phoneme_end|>見えます。 The pitch digits disambiguate homographs like 端, 箸 and 橋, which share the same sounds.

Keep punctuation outside the tags, and tag short spans rather than whole paragraphs.

For rhythm, write natural pause words like “um” and “uh” straight into the text, or use parenthesis effects: (break) for a short pause, (long-break) for a longer one, plus (breath), (laugh), (cough), (lip-smacking) and (sigh). The last four are experimental and may need repeating to land.

I am, um, an (break) <|phoneme_start|>EH1 N JH AH0 N IH1 R<|phoneme_end|>.

What creatives use it for

  • Narration. Audiobooks, documentaries and explainers, with emotion tags doing the vocal direction a session with a voice actor would.
  • Character dialogue. Distinct deliveries for animation and game scripts; layer emotions per line and keep takes consistent across revisions.
  • Video voiceover. Scratch or final VO for edits, ads and social cuts, regenerated in seconds every time the script changes.
  • Localization. The same script and the same voice across 13 languages.
  • Your voice, at scale. Clone your voice once and narrate everything in it without booking yourself into a booth.

FAQ

How much text can I convert at once? Up to 50,000 characters per run, enough for very long scripts and full chapters.

Which model should I pick? S2.1 Pro, the default. It’s the latest and highest quality. S2 Pro and S1 are previous generations, useful when you need to match audio you generated with them.

Can it speak in my voice? Yes. Clone your voice with the Voice Cloner first; the cloned voice appears under My presets and you generate speech from there. The additional settings also accept a Fish Audio voice ID directly.

How do I control emotion? With inline tags in square brackets: [happy], [whispering], [sarcastic] and 60+ more, including free-form descriptions like [warm and happy]. See the full reference above.

Which languages does it support? 13: English, Chinese, Japanese, German, French, Spanish, Korean, Arabic, Russian, Dutch, Italian, Polish and Portuguese. Emotion tags work in all of them.

What audio formats can it output? MP3 (at 64, 128 or 192 kbps), WAV, or Opus. MP3 at 128 kbps is the default.

Can I change the speaking speed? Yes, from 0.5x to 2x in the additional settings.

Got a script waiting? Give it a voice on mitte.ai.

Related tools

Auto-subtitles Brefnet Background Remover Bria Background Remover Crystal Image Upscaler Crystal Video Upscaler Demucs Track Splitter

All AI models & tools

See pricing and start creating on Mitte

Mitte AI Models & Tools Presets Pricing Enterprise Creators Careers About Blog