Flux 3

Create videos up to 20 seconds in one generation, with sound built in.

FLUX 3 AI Video Generator: Text, Images & Keyframes to Video with Native Audio

Create videos up to 20 seconds in one generation, with sound built in. FLUX 3 is Black Forest Labs’ video model, and Mitte gives you a simple interface for all of it: type a prompt, start from an image, pin keyframes, or continue a video you already have.

Every clip comes with native audio — ambience, sound effects, and spoken dialogue. Dialogue works in multiple languages with strong lipsync, and the model is especially good at faces and at matching sounds to what happens on screen.

What can you do with it?

How to use it

  1. Pick a mode. Keyframes, First / Last Frame, Text to Video, or Extend Video.
  2. Add your inputs. Drop in your images or video depending on the mode — or nothing at all for text to video.
  3. Describe the shot. What happens, how the camera moves, what you hear. Write dialogue lines directly in the prompt.
  4. Set duration, aspect ratio, and resolution. 5 to 20 seconds, seven aspect ratios from 21:9 to 9:16, at 720p or 1080p.
  5. Run it. You get an MP4 with audio, ready to download, share, or feed into your next step on Mitte.

Modes

Settings

Tips

All AI models & tools

See pricing and start creating on Mitte