Kling O3

Generates **3 to 15 seconds of video with sound in one pass**, at 1080p in Pro or in full 4K.

Kling O3: Video with Native Sound from Text, Images, and a Reference Clip

Kling O3 is Kuaishou’s newest video model. It generates 3 to 15 seconds of video with sound in one pass, at 1080p in Pro or in full 4K. Give it a text prompt, a first frame, a set of reference images, or a reference video, and tag each file right in the prompt: @Image1 for a subject, @Video1 for the clip whose motion you want.

The reference video is the headline. Attach a clip and O3 follows its motion, framing, and timing while you swap in new subjects, a new setting, or a new style from your images. The result is a new video of the length you choose, with the clip’s sound kept or dropped.

What can you do with it?

How to use it

  1. Pick your mode. Omni Mode takes a prompt plus reference images and at most one reference video. Frame Mode takes a first frame, an optional last frame, and a motion prompt.
  2. Attach your references. Up to 7 images on their own, or one video plus up to 4 images. Images should be at least 300 px on each side. The video should be 3 to 15 seconds, MP4 or MOV, 720p or better.
  3. Write the prompt. Call out each reference by its tag and say what it contributes: “@Image1 is the dancer, @Video1 is the choreography.” Say what should stay and what should change.
  4. Set model, duration, and ratio. Standard, Pro, or 4K. 3 to 15 seconds. 16:9, 9:16, 1:1, or Auto to follow your reference video.

What you get

One video, 3 to 15 seconds, as mp4. Audio is generated with the video, or taken from your reference clip when you attach one and keep its sound. In frame mode the output ratio follows your first frame.

Tips

All AI models & tools

See pricing and start creating on Mitte