Wan 3.0

Wan 3.0 is Alibaba's next-generation video model. It generates **up to 30 seconds of 1080p video in a single pass**, with audio baked into the same generation. One continuous shot, no stitched clips.

Wan 3.0: 30 Seconds of 1080p Video, With Sound, in One Pass

Wan 3.0 is Alibaba’s next-generation video model. It generates up to 30 seconds of 1080p video in a single pass, with audio baked into the same generation. One continuous shot, no stitched clips.

Give it a text prompt, a first frame, or a whole pile of references: up to 10 images, 5 video clips, and 5 audio tracks in one request. Turn on Thinking and it can even build a video from a public webpage.

What can you do with it?

How to use it

  1. Pick your mode. Reference Mode takes a prompt plus optional reference media. Frame Mode takes a first frame, an optional last frame, and a motion prompt.
  2. Attach your references. Up to 10 images for subjects and style, 5 video clips for motion (15 seconds combined), 5 audio tracks for sound (15 seconds combined).
  3. Write the prompt. Address each reference by position (Image 1, Video 2) and say what it contributes. Be concrete about camera, light, and pacing.
  4. Set resolution, duration, and ratio. 480p, 720p, or 1080p. 2 to 30 seconds. Any ratio from 16:9 to 9:16, or adaptive to let the model decide.

What you get

One video, 2 to 30 seconds, as mp4 with audio (or without, if you switch it off). Adaptive ratio follows your prompt and references; otherwise you get exactly the ratio you picked.

Tips

All AI models & tools

See pricing and start creating on Mitte