Text to Speech
Turn a script into natural, expressive speech — direct emotion, tone, and delivery with simple inline tags. Runs on Fish Audio's S2 models.
AI Text to Speech
Turn text into natural, expressive speech — with inline tags to control emotion, tone, and delivery.
Paste your script, choose a voice, and generate. This tool reads it aloud in a natural voice and lands the result in your library as an audio asset, ready to drop into an edit. It runs on Fish Audio’s S2 models, so you’re not stuck with a flat, robotic read — you direct the performance from the text itself.
How to convert text to speech
- Paste or type your script — up to 50,000 characters.
- Pick a voice: use the default, or speak in a voice you’ve cloned with the Voice Cloner or invented with the Voice Designer.
- Generate. The audio lands in your library, ready to use in a video, podcast, or app.
Emotion control
Direct the performance right from your script with inline tags in square brackets:
-
Wrap an emotion in brackets before the words it should affect:
[happy],[sad],[excited],[angry],[whispering],[shouting]. -
Combine tags for nuance:
[sad][whispering] I miss you so much. -
Adjust intensity with modifiers:
[slightly sad],[very excited],[extremely angry]. - Keep one primary emotion per sentence, placed at the start of the words it controls, for the most natural read.
There are 64+ emotion and delivery tags in total, plus phoneme tags for exact pronunciation of names and technical terms. See the full tag reference on the node page for the complete list and advanced usage.
Features
- 64+ emotion and delivery tags — happy, sad, excited, whispering, shouting, and more
- Up to 50,000 characters per generation
- 13 languages supported
- Works with your own cloned or designed voices
- Exact pronunciation control via phoneme tags for names and technical terms
FAQ
How do I add emotion to an AI voice?
Wrap an emotion word in square brackets right before the text it should affect, like [excited] We won!. You can combine tags, like [sad][whispering], and adjust intensity with words like “slightly” or “very”. See the full tag reference on the node page for the complete list.
How do I make text-to-speech sound more natural?
Use one primary emotion per sentence, keep the tag right next to the words it controls, and avoid stacking conflicting emotions in a single sentence. Natural-sounding delivery comes from matching the tag to the actual context of the line, not from piling on tags.
What’s the difference between this and the Voice Cloner or Voice Designer?
This tool reads any text aloud in a chosen voice. The Voice Cloner and Voice Designer are how you get that voice in the first place — cloning one from a real recording, or inventing one from a description. Once you have a voice, use this tool to make it speak.
How many languages does it support?
13: English, Chinese, Japanese, German, French, Spanish, Korean, Arabic, Russian, Dutch, Italian, Polish, and Portuguese.