Wan 3.0
Wan 3.0 is Alibaba's next-generation video model. It generates **up to 30 seconds of 1080p video in a single pass**, with audio baked into the same generation. One continuous shot, no stitched clips.
Wan 3.0: 30 Seconds of 1080p Video, With Sound, in One Pass
Wan 3.0 is Alibaba’s next-generation video model. It generates up to 30 seconds of 1080p video in a single pass, with audio baked into the same generation. One continuous shot, no stitched clips.
Give it a text prompt, a first frame, or a whole pile of references: up to 10 images, 5 video clips, and 5 audio tracks in one request. Turn on Thinking and it can even build a video from a public webpage.
What can you do with it?
- Write a scene and get a finished shot. Prompt it and Wan generates picture and sound together: dialogue scenes, ambient noise, music.
- Animate a still image. Your image becomes the first frame and the video grows out of it.
- Control the start and end. Frame mode takes a first and a last frame and generates the motion between them.
- Keep subjects consistent. Attach reference images of a person, character, or product and call them out in the prompt: “the subject in Image 1 walks past Video 1.”
- Borrow motion and sound. Reference clips steer the movement, reference audio steers the soundtrack.
- Turn a webpage into a video. Paste a public URL (with Thinking on) and Wan reads the page and builds a video from it.
- Direct real movement. Wan holds up under full-body action: dancers, athletes, fight choreography, without limbs melting mid-move.
How to use it
- Pick your mode. Reference Mode takes a prompt plus optional reference media. Frame Mode takes a first frame, an optional last frame, and a motion prompt.
- Attach your references. Up to 10 images for subjects and style, 5 video clips for motion (15 seconds combined), 5 audio tracks for sound (15 seconds combined).
- Write the prompt. Address each reference by position (Image 1, Video 2) and say what it contributes. Be concrete about camera, light, and pacing.
- Set resolution, duration, and ratio. 480p, 720p, or 1080p. 2 to 30 seconds. Any ratio from 16:9 to 9:16, or adaptive to let the model decide.
What you get
One video, 2 to 30 seconds, as mp4 with audio (or without, if you switch it off). Adaptive ratio follows your prompt and references; otherwise you get exactly the ratio you picked.
Tips
- Draft at 480p, finish at 1080p. 480p costs a quarter of 1080p per second, so iterate cheap and rerender the keeper.
- Name the role of every reference. “Image 1 is the hero, Video 1 is the camera move” beats attaching files and hoping.
- Turn on Thinking for complex scenes. It reasons about composition and motion before generating. Slower, but it pays off on multi-subject shots and webpage inputs.
- Leave Prompt Expansion on. Switching it off shaves 20 to 60 seconds of wait but usually costs you quality.
- Long doesn’t have to mean slow. A single 30-second pass holds one continuous action better than 6 stitched 5-second clips ever will. Write the prompt as one scene with beats, and let it play out.