Wan 3.0
AI Video Generator

Turn a prompt, a photo or a set of references into a 2-30 second video with native sound in 480p, 720p or 1080p, using the Wan 3.0 AI video generator.

The Wan 3.0 AI video generator produces up to 30 seconds of picture and sound in a single generation, keeping characters, props and camera direction consistent from the first shot to the last.

An animated woman in a pink jacket clutches a glowing orb as a man in a blue hoodie grabs for it on a neon city street.

Multi-Reference Action in One Take One image sets the neon city and one each fixes the two characters. Wan 3.0 keeps both faces and outfits intact through a 30-second chase on foot and by bike.

A crane claw lowers a model in a light-blue jacket and tan shorts onto a pastel studio pedestal.

Vertical Fashion Films Frame for 9:16 and a full 30 seconds holds one model, one outfit and one clean studio set, ready for Reels and Shorts.

A giant sea creature rises out of a night storm beside a small fishing boat.

Scale, Weather and Sound Storm waves, lightning and a creature rising from the sea hold their scale against a small fishing boat, with the sound generated alongside the picture.

A silver and black compact camera floats against a pale backdrop.

Product Films from One Brief Describe the sequence from an aperture macro to the hero shot, and Wan 3.0 builds a 15-second product spot with clean light on metal and glass.

Key Features of Wan 3.0

Build complete scenes with the Wan 3.0 AI video generator, from 30-second takes to videos guided by the images, clips and audio you upload.

30 Seconds in a Single Generation

Choose any length from 2 to 30 seconds, enough to follow a roll of film from the steel reel to the drying line. Wan 3.0 carries the light, focus and pacing through every step.
Generate a 30-second video
A developing negative shows a smiling girl under amber darkroom light.

20 References in One Generation

Combine 10 images, 5 videos and 5 audio clips in one generation. Upload a model's portrait and a garment photo, and the face, embroidery and cut stay intact from the fabric close-up to the runway walk.
Create from references
A model with platinum hair walks a dark runway in a black qipao embroidered with peonies.

Native Sound and Score

Dialogue, effects and music are generated with the picture. Name the instruments, the impacts and the moment each should land, and the soundtrack follows the action on screen.
Create a video with sound
A dancer in a jeweled headdress faces a fiery-haired giant inside a dim cave temple.

Lifelike Faces and Physics

Faces come out varied and natural, with micro-expressions that match the body language, and water, cloth and hair move with real weight. Start from a still frame and the scene keeps its look once the action begins.
Animate a photo
A soaked young woman pushes through seawater flooding a wood-paneled ship corridor.

Wan 3.0 Creative Controls

The Wan 3.0 AI video generator follows direction on references, timing, framing and sound.

References by Number

Refer to uploads as Image 1, Video 1 and Audio 1, counted in the order you add them. Say what each one is for, such as which image sets the location and which one is the lead character.

Timed Prompts for Every Beat

Break a long take into timed beats, such as 0-4s for the fly-through over the city and 4-8s for the first face-off. Wan 3.0 moves through the beats in order.

First and Last Frame

Upload an opening image and a closing image, and Wan 3.0 fills in the motion between them. Use it when a product reveal or a transition has to land on an exact shot.

Style References

Add a style frame and write that it sets the look, for example "Refer to Image 1 for the aesthetic style." A clay, anime or film look then carries across the whole video while your subject stays the same.

Sound Written into the Prompt

Write dialogue in quotes and describe ambience, effects and music in plain words. Switch sound off when you plan to score the video yourself.

Length, Resolution and Ratio

Choose 2-30 seconds in 480p, 720p or 1080p. Frame 16:9 for film, 9:16 for Reels and Shorts, 1:1, 4:3 or 3:4 for feeds, or Auto to follow the shape of your input.

Made with Wan 3.0

The Wan 3.0 AI video generator follows long, detailed direction, including shot lists, camera moves, reference roles and sound design.

Rows of young Shaolin novice monks in grey robes train in a temple courtyard.

Shaolin training drill

In the stone courtyard behind Shaolin Temple, rows of young novice monks repeat high kicks, side kicks and spinning kicks in perfect sync on the instructor's command. Cut to one boy holding a one-legged stance, palms pressed together, as the instructor's staff corrects his elbow. Then senior disciples drill on wooden dummies. Only wooden impacts, shouts and echoes.

A couple in period evening wear waltzes across a marble ballroom as guests look on.

19th-century ballroom waltz

Character A wears a black tailcoat, white shirt, bow tie and beige waistcoat. Character B wears a beige puff-sleeve gown with black lace trim. Under crystal chandeliers the guests fall quiet as a string waltz begins. B walks in, A crosses the room, bows and offers his hand, and the crowd steps back as the two begin to turn. Warm golden tones.

A clay fox noses a red apple at a miniature fruit stall.

Clay fox at the market

Refer to Image 1 for the aesthetic style. A clay-textured little fox runs down an autumn hillside, crosses a tiny stone bridge over a resin-like stream and stops at a market fruit stall. It paws a red apple loose, carries it off in its mouth and eats it on a low stone wall at sunset. Miniature scene, shallow depth of field.

A young girl in a white summer dress runs along a sunny beach boardwalk chasing bubbles.

Seaside bubble commercial

A light, dreamy live-action commercial. A six-year-old girl in a summer dress stands on a seaside boardwalk, dips a small wand into a bottle of bubble solution and blows. The sea breeze pushes the bubbles sideways across the frame, and she reaches out to pop them, laughing. Ocean-blue tones, golden backlight, rainbow reflections on every bubble.

A tiny winged fairy hovers beside a large teddy bear in a child's toy room.

Dragonfly fairy in a toy room

One continuous 10-second shot, low over a toy-room carpet. A tiny fairy with dragonfly wings and a pink skirt hovers, then darts across the room in sharp bursts, stopping dead each time: past a teddy bear, behind a red toy truck, up the towers of a block castle. The castle collapses and she dives through it into deep-blue water. Style reference: Image 2.

A hooded fighter and an armored soldier grapple over a glowing blue sword in the rain.

Rain-soaked anime duel

Stylized 3D animation at night in heavy rain. A hooded fighter in a dark cloak and a soldier with armored sleeves trade blows on a wet rooftop lit by blue holograms. Low angle as a glowing blade sweeps past the soldier's face, then a close-up as they lock hands over the hilt. Rain hiss and metal impacts.

Wan 3.0 Use Cases

Put the Wan 3.0 AI video generator to work on films, ads and social content where a scene has to run long and stay consistent.

Film Scenes and Previs

Film Scenes and Previs

Write the scene as a timed shot list with lens, pacing and blocking, and Wan 3.0 stages it as one continuous take. A battle, a chase or a reveal can be tested before the first day on set.

Product Ads and Brand Films

Product Ads and Brand Films

Upload product photos as references and list the shots, from a macro of the liquid to the final hero frame. Logos, materials and proportions stay fixed from shot to shot, ready for product pages and paid campaigns.

Animated Shorts

Animated Shorts

Add your character as Image 1 and describe the adventure in timed beats. A 20 or 30-second generation holds a full arc, from the setup to the escape, in one 3D, clay or anime style.

Music and Dance Videos

Music and Dance Videos

Give each performer a character reference and describe the formations as the track builds. Faces, outfits and positions hold through fast cuts, ready for YouTube and TikTok.

Travel and Destination Films

Travel and Destination Films

Describe the streets, landmarks and food of a place, and Wan 3.0 returns a walk-through in natural light. Add hand-drawn characters to the live-action street for a lighter tone, without a location shoot.

Food and Recipe Videos

Food and Recipe Videos

Split a recipe into timed shots, such as 1-3s for the egg dropping into flour and 3-6s for the kneading. Macro detail on dough, flour and hands holds up for recipe reels and restaurant ads.

How to generate

Three steps in the Wan 3.0 AI video generator

  1. Add an image

    Step 01

    Add an image

    Upload a photo to animate, or add image, video and audio references.

  2. Describe the scene

    Step 02

    Describe the scene

    Write the action, camera movement and sound. Put dialogue in quotes.

  3. Generate and download

    Step 03

    Generate and download

    Set the length, resolution and aspect ratio, then generate and download.

A woman in a long red cape stands in shallow water as black horses gallop past.

Try the Wan 3.0 AI Video Generator

Start with a prompt, a photo or a handful of references. Generate a 30-second scene with sound, compare takes and download the one you want in 1080p.

Pricing

Upgrade or cancel anytime. Choose a plan for clean output, fast generation, full model access and priority support.

Monthly

Annual

36% OFF

Wan 3.0 AI Video Generator FAQs

Quick answers about Wan 3.0, plus the plans, credits and billing behind it on ImgVid.

Wan 3.0 is the newest generation of Alibaba's Wan video models, built by the Tongyi Wanxiang team and released in public beta on August 6, 2026. It generates 2 to 30 seconds of video with native sound in a single pass and can follow images, video clips and audio as references. On ImgVid we've set it up for text to video, image to video, reference to video and first and last frame to video, all in the same workspace.

Three things stand out. A single video runs to 30 seconds, twice the 15 seconds of Wan 2.7. On ImgVid the reference limit grows from 4 images and 3 videos to 10 images and 5 videos, alongside 5 audio clips. And faces look more varied and lifelike, while characters, props, spaces and style hold steadier across shots. Wan 3.0 also adds 480p, which makes drafts cheaper than on Wan 2.7. Alibaba's own API can also read documents and web pages and edit existing videos, but those options aren't part of ImgVid's setup yet.

Any length from 2 to 30 seconds, in 480p, 720p or 1080p. Aspect ratios cover 16:9, 9:16, 1:1, 4:3 and 3:4, plus Auto, which follows the shape of your input image. Our tip: draft at 480p and save 1080p for the take you plan to publish.

Yes. Dialogue, sound effects and background music are generated with the picture, so there's no separate dubbing pass. Put spoken lines in quotes, name the speaker, then describe the ambience and music you want. Alibaba notes that audio texture and on-screen text are still being refined, so keep signs and captions short and check dialogue-heavy takes before you publish. Prefer a silent clip? Just switch the sound off.

Up to 20 per video: 10 images, 5 video clips and 5 audio files. The prompt refers to them as Image 1, Video 1 and Audio 1 in the order you upload them, so say what each one is for, for example which image is the lead character and which sets the location. First and last frame video runs as a separate mode, so pick one approach per generation. Only upload material you have the rights to use.

No. Earlier versions such as Wan 2.1 and Wan 2.2 were released with open weights, but as of September 2026 Alibaba offers Wan 3.0 only as a hosted model, and its official Hugging Face and GitHub pages carry no Wan 3.0 weights. On ImgVid you use it in the browser, with no setup and no GPU of your own.

Wan 3.0 is included in the Plus and Premium plans, which unlock every model on ImgVid, including the earlier Wan versions and every Seedance model. Free and Starter include Seedance 2.0 Fast and Seedance 2.0 Mini along with a set of image models, a good way to test ideas before you upgrade.

It depends on length and resolution. Per second, Wan 3.0 uses about 3 credits at 480p, 6 at 720p and 13 at 1080p, so a 5-second 720p clip costs 32 credits and a 30-second 1080p video costs 378. The exact number appears in the generator before you confirm, and if a generation fails the credits go straight back to your balance.

Not on the Free plan itself, since Wan 3.0 comes with Plus and Premium. What you can do for free is sign up with no credit card, get a monthly batch of credits and try Seedance 2.0 Mini or image models such as GPT Image 2 and Seedream 5 Lite to see how the workspace feels. Free exports carry a watermark. Any paid plan removes it, and everything you generate on a paid plan is yours to use commercially.

Plus comes with 2,000 credits a month, enough for about 62 five-second 720p clips or 15 ten-second 1080p videos with Wan 3.0. Premium comes with 6,000, three times as much. Subscription credits stay valid for two months, so an unused balance carries into the next cycle. If you run short mid-month, a credit pack tops you up instantly and never expires.

Yes to both. Cancel from the Membership tab in your account settings and you keep every benefit until the end of the period you've paid for, after which the account returns to the Free plan. Refunds are available on monthly and annual orders, based on how much of the plan you've used, and our team handles requests within 24 hours.