Whisper

Upload a recording and get the full transcript back.

Audio Transcription: Turn Speech into Text with Word-Level Timestamps

Upload a recording and get the full transcript back, plus the exact start and end time of every single word. Interviews, podcasts, voice memos, lectures: if someone’s talking in it, this turns it into text you can search, quote, subtitle, or feed into your next step.

It runs on OpenAI’s Whisper, trained on 680,000 hours of speech across roughly 100 languages. You don’t need to tell it what language the audio is in; it figures that out on its own.

What can you do with it?

How to use it

  1. Upload your recording. Any common audio format works (mp3, wav, m4a, flac), up to 25MB.
  2. Optionally set the hints. Pick the spoken language if you know it, and add names or jargon in the hint field so they come out spelled right (“Mitte”, “Kubernetes”, “Dr. Okonkwo”).
  3. Run it. A typical recording transcribes in well under a minute.

Inside a workflow you can skip the upload entirely: the transcription step picks up the audio produced by the step before it, like a voice generator or a stem splitter’s vocal track.

What you get

Structured JSON, ready for downstream steps:

Tips

All AI models & tools

See pricing and start creating on Mitte