Multimodal reference control
Guide generation with up to 10 images, 5 video clips, 5 audio clips, or one document or public web page.
Understands your media, documents, and web pages, giving video generation richer creative context.
Wan 3.0 AI video generator
Let images define the look, video guide motion, audio carry the rhythm, and documents and web pages add background—together giving video generation richer, more complete context.
Guide generation with up to 10 images, 5 video clips, 5 audio clips, or one document or public web page.
Choose 480P, 720P, or 1080P and use Adaptive, 16:9, 4:3, 1:1, 3:4, or 9:16 aspect ratios.
Write prompts up to 20,000 characters, use start and end frames, enable synchronized audio, and reuse a seed for repeatable settings.
Combine product images and documents to turn features, positioning, and visual direction into a concise promotional video.
Use audio to guide rhythm and mood, with images or video defining performers, scenes, and motion.
Build characters, locations, camera language, and story beats from multiple visual and text references.
Turn a supported file or public web page into a product explainer, content summary, or presentation video.
Define characters, products, scenes, and visual style.
Guide motion, camera movement, and pacing.
Provide music, voice, rhythm, and atmosphere.
Add product facts, scripts, and story background.
Supply public page content and its surrounding context.
Upload media or a document, add a public web page, or choose start and end frames.
Write the desired result and use Image1, Video1, or Audio1 to point to specific media.
Select generation mode, duration, resolution, aspect ratio, and audio, then start generation.
Faster generation for shorter waits