Flux 3
Create videos up to 20 seconds in one generation, with sound built in.
FLUX 3 AI Video Generator: Text, Images & Keyframes to Video with Native Audio
Create videos up to 20 seconds in one generation, with sound built in. FLUX 3 is Black Forest Labs’ video model, and Mitte gives you a simple interface for all of it: type a prompt, start from an image, pin keyframes, or continue a video you already have.
Every clip comes with native audio — ambience, sound effects, and spoken dialogue. Dialogue works in multiple languages with strong lipsync, and the model is especially good at faces and at matching sounds to what happens on screen.
What can you do with it?
- Turn an idea into a finished clip. Describe the shot and get video with sound, dialogue included.
- Bring a still image to life. Your image becomes the first frame and the scene moves from there.
- Control the story with keyframes. Pin up to 10 images to exact moments in the timeline and let the model fill in the motion between them.
- Morph between two images. Set a first and last frame and get a smooth transition connecting them.
- Make longer videos piece by piece. Extend a clip and FLUX 3 carries the framing and scene logic forward, so chained clips feel like one continuous video.
- Shoot dialogue scenes without a shoot. Characters speak your lines, lipsynced, in the language you write them in.
How to use it
- Pick a mode. Keyframes, First / Last Frame, Text to Video, or Extend Video.
- Add your inputs. Drop in your images or video depending on the mode — or nothing at all for text to video.
- Describe the shot. What happens, how the camera moves, what you hear. Write dialogue lines directly in the prompt.
- Set duration, aspect ratio, and resolution. 5 to 20 seconds, seven aspect ratios from 21:9 to 9:16, at 720p or 1080p.
- Run it. You get an MP4 with audio, ready to download, share, or feed into your next step on Mitte.
Modes
- Keyframes: pin 1–10 images to exact frame positions; the model generates the motion between them.
- First / Last Frame: set the start and end images. Leave the last frame empty to simply animate from the first.
- Text to Video: no images needed — the prompt is the whole brief.
- Extend Video: continue an existing clip (MP4, under 15 seconds) from its final frames.
Settings
- Duration: 5–20 seconds per generation.
- Aspect Ratio: 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, 9:16 — or Auto to match your images.
- Resolution: 720p for speed and cost, 1080p for quality.
- Audio: on by default; switch it off if you’re scoring the video yourself.
Tips
- Write the soundtrack into the prompt. Mention ambience, effects, and dialogue lines — audio is generated with the video, not added after.
- Use keyframes for story beats. Three or four pinned images give you a storyboard the model animates, far more control than a prompt alone.
- Chain clips with Extend Video. Generate 20 seconds, extend from the ending, repeat — framing and scene logic carry over, so the result cuts together as one long shot.
- Reuse character images across scenes. Starting new clips from images of the same character keeps them consistent from shot to shot.
- Iterate at 720p, finish at 1080p. Try versions cheaply, then rerun the keeper at full quality.