Wan 3.0 Video

Wan 3.0 Video

Wan 3.0 Video is Alibaba's multimodal video model for text-to-video, image-to-video, reference-to-video, and video editing with image, video, and audio references.

About

Wan 3.0 Video

Wan 3.0 Video is Alibaba's multimodal video model for text-to-video, image-to-video, reference-to-video, and video editing with image, video, and audio references.

Wan 3.0 Video
How it works

Create with Wan 3.0 Video

Wan 3.0 Video is Alibaba's multimodal video model for text-to-video, image-to-video, reference-to-video, and video editing with image, video, and audio references.

01

Describe the shot

Write the subject, action, camera movement, lighting, and sound you want in one clear prompt.

02

Add references

Use an image, video, or audio reference when the subject, motion, style, or rhythm needs tighter control.

03

Generate and refine

Choose the duration, aspect ratio, and resolution, then iterate on the prompt until the shot is ready.

Duration

5 / 10 / 15 / 20 / 25 / 30

Resolution

480p / 720p / 1080p

Aspect Ratio

adaptive / 16:9 / 4:3 / 1:1 / 3:4 / 9:16

Reference Image

Max 10 images for video generation

Wan 3.0 Video — Settings

Wan 3.0 Video is Alibaba's multimodal video model for text-to-video, image-to-video, reference-to-video, and video editing with image, video, and audio references.

Multimodal references

Suitable Scenes: multimodal video creation, product showcases, advertising clips, and creative short films with image, video, or audio references.

Native audio support

Prompt tips: Describe the subject, motion, camera direction, and the role of each reference. Choose duration and resolution for the intended output.

480p to 1080p

Product videos,Brand advertising,Short films,Social content

Examples & use cases

What you can make with Wan 3.0 Video

Explore practical ideas for camera direction, reference control, character consistency, native audio, art direction, and precise prompt timing.

01

Camera & motion

Camera moves you can direct

Describe a tracking, orbit, tilt, or push-in shot and keep the movement tied to the subject, background, and pacing.

View example prompt

A product slowly rotates on a studio table under soft light with crisp reflections.

02

First & last frame

Two references become one shot

Use a start frame and an optional end frame when the opening and closing composition need to stay intentional across the transition.

View example prompt

A character walks through a night city while following the motion from the reference video.

03

Character consistency

The same character, shot after shot

Keep the character identity, clothing, proportions, and visual traits stable while the location, lighting, or camera changes.

View example prompt

Create a short product commercial using reference images, audio pacing, and a cinematic camera move.

04

Native audio

Sound generated with the picture

Describe dialogue, ambience, foley, music, or rhythm together with the shot so the audio direction follows the visual action.

View example prompt

A short cinematic shot with synchronized ambience, foley, and a clear audio rhythm that matches the movement on screen.

05

Art direction

One look held across every shot

Give the model a visual language—palette, lighting, linework, or material—and carry it consistently through the whole clip.

View example prompt

A sequence with a consistent palette, lighting language, texture, and art direction from the first frame to the last.

06

Prompt adherence

It hits the beats in the order you write them

Put the subject, action, camera, timing, and constraints in a clear order; precise prompts make complex shots easier to control.

View example prompt

A precise shot brief with ordered actions, camera direction, timing, lighting, and output constraints.

FAQ

Wan 3.0 Video FAQ

Wan 3.0 Video is Alibaba's multimodal video model for text-to-video, image-to-video, reference-to-video, and video editing with image, video, and audio references.

Explore More AI Video Models

Wan 3.0 Video

Wan 3.0 Video is Alibaba's multimodal video model for text-to-video, image-to-video, reference-to-video, and video editing with image, video, and audio references.