Model · Alibaba Wan 3.0

Wan 3.0 — 30 seconds of video in one run

Alibaba's video model, and the one thing it does that Google's models here don't: length. A single generation gives you 4 to 30 seconds with sound already in it, at 480p, 720p or 1080p, horizontal or vertical. You pay by the second — $0.05, $0.10 or $0.20 depending on resolution — so a short draft costs $0.20 and the longest, sharpest clip on the platform costs $6.00. Crypto or card, no subscription.

A real Wan 3.0 generation on BananaBanana: one prompt, ten seconds at 720p, sound rendered in the same pass. That clip cost $1.00. Turn the audio off and the price does not change.
01What it is

What is Wan 3.0?

Wan 3.0 is Alibaba's video generation model, released on Alibaba Cloud Model Studio in August 2026, and it renders clips of 4 to 30 seconds with audio generated in the same pass as the picture. Here it reads a prompt, an optional starting frame and closing frame, or up to 10 reference images, and returns one finished clip at 480p, 720p or 1080p. We charge exactly what Alibaba lists for the model — $0.05, $0.10 and $0.20 per second by resolution — rather than marking it up, which is the same policy we apply to Google's models.

Who actually needs 30 seconds?

Anyone who is tired of stitching. Most video models cap out around eight or ten seconds, so a half-minute sequence means three or four generations, matching them by hand, and living with the seams where the model changed its mind about the lighting. Wan renders the whole thing as one take. For product loops, ambient backgrounds, process shots and anything where the camera just keeps moving, that removes the editing step entirely. Vertical works the same way — 9:16 at any of the three resolutions, which is what Shorts and Reels want.

Vertical 9:16 at 480p, six seconds, sound included. $0.05 a second means this one cost thirty cents.
02What it does well

Six things worth knowing before you spend

Up to 30 seconds, one take

4, 6, 8, 10, 15, 20 or 30 seconds. No stitching, no drift between segments. The next longest model here reaches fifteen; Google's stop at eight or ten.

Sound that costs nothing

Audio is generated with the picture and is on by default. Switching it off is a toggle, and the price stays the same either way — unlike Veo, where native audio is a paid extra.

First and last frame

Give it a starting image to animate, and optionally a closing image to land on. Same picture in both slots gives you a clean loop.

Up to 10 reference images

References keep a subject or a style consistent across runs. They are an alternative to frames, not an addition — the model takes one or the other.

Seeds

Fix the seed and the same prompt lands close to the same shot. Useful when you are tuning one word at a time and want to see what that word did.

Three price tiers

480p for drafts at $0.05 a second, 720p for publishing at $0.10, 1080p at $0.20. Draft cheap, render the keeper once.

Twenty unbroken seconds from a single generation at 480p — clay centred, wall pulled, rim opened, no cuts anywhere. It cost a dollar. The hands stay coherent the whole way through, which is not something I expected from a twenty-second take.
03Honest limits

What Wan 3.0 will not do

Editing here is not a conversation. With Gemini Omni Flash you can say "make the light warmer" and the model adjusts the scene it remembers. Wan can re-shoot a clip you hand it — up to 15 seconds of source, with source and result together inside 30 seconds — but every request starts from zero, and the seconds of the source are billed along with the seconds of the result. It also renders one clip per run — no four-variant grids to pick from. There is no negative prompt, so anything you do not want has to be written out of the scene rather than banned from it. Aspect ratios are 16:9 and 9:16 only, and 1080p is the ceiling. And starting frames and reference images are mutually exclusive: pick frames or pick references, not both.

Alibaba is candid about where the model is still growing: its own launch notes say audio texture and on-screen text rendering are not yet where they want them, and that roughly matches what we see — the ambience lands, spoken lines came out better than we expected, but it is not a track you can shape in detail. Long takes carry their own risk. Our twenty-second pottery clip above held together better than I expected, but thirty seconds is a lot of frames to keep consistent, so budget a retry for anything past fifteen. One thing the model does upstream that we have not wired in: it accepts documents (PPT, PDF, spreadsheets) as input. And its "extend" is a next scene with the same cast and set, not a frame-accurate continuation. If you want audio you can steer sentence by sentence, Gemini Omni Flash is the better tool; if you need 4K, which Wan does not reach, Veo 3.1 is the alternative.

04Field notes

What production work taught us

These notes come from real ad production on Wan 3.0 in September 2026: product clips, clean-up passes over filmed footage, remakes built from a first frame. None of it is in the documentation, and all of it cost money to find out.

Block order decides the outcome

Whatever comes first in the prompt gets the attention, and a long product description pulls it all. In our runs a clean-up instruction sitting seventh in the prompt was skipped; the same words moved to the top as "TASK 1" were carried out. Put the thing you need most first.

Negatives work as a drawing list

"No ring, no watch, no ink" gave us a new gold ring and a bolder tattoo. In our runs, naming a thing read as asking for it. Describe what should be there instead: "each finger is plain skin from base to tip" did what three negatives could not.

It fills empty hands

In our shots where the product was not supposed to appear, Wan swapped whatever the hand was holding for it — a comb turned into a trimmer. Name every tool, say out loud that a frame without the product is fine, and write "never more than one" instead of "exactly one". Safer still: render those shots as a separate segment whose prompt never describes the product.

A first frame holds the object

From one starting image and no references, Wan kept a product intact for ten seconds: buttons in place, nothing extra in the scene. In our tests on the same frame, other models grew a stray cable and foreign objects. If the exact product matters, make the frame first and let Wan animate it in one run.

Clean-up comes back frame for frame

Editing a filmed clip to remove tattoos, a watch and bracelets returned footage that matched the source shot for shot, with clean skin. Pieces of up to about eight seconds were the reliable size for us; longer sources are better cut into parts.

One long run beats a chain

Extend on Wan means "the next scene", not a frame-accurate continuation. When you need fifteen, twenty or thirty seconds, order them as one generation instead of chaining short ones.

Durations are whole seconds

The result is trimmed to the whole number of seconds you ordered, at 30 fps: 94 frames in came back as 90. Cutting exactly on a shot boundary therefore did not work for us — plan your segments so that the joins fall on whole seconds.

A reference photo brings its wardrobe

In our runs Wan copied clothes and jewellery from a photo of a person even when the prompt said "face and hair only". Crop the reference down to what you actually want carried over.

Room for a shot-by-shot brief

The prompt takes up to 20,000 characters, so there is space for numbered tasks, a lock on every tool and a description of each shot. Wan rewards that structure: the more exactly a block says what is in frame, the less it improvises.

Defects to plan around. Sound is normally there, and the model delivers spoken lines decently — not only English ones. One project went the other way: in edit passes and in a text-to-video run with a complex prompt, the sound of a trimmer motor came back quiet, around −45 dB — audible, but weak. In our other runs the track was where it should be, so that is one situation rather than a trait of the model; if a thin track would hurt, check the level and take the sound from your source when you need to. Small digits on a device display are drawn better than by the competition but not reliably — we got 89 where the frame said 84, and with a horizontal grip the digits came out rotated. The shape of a product can soften in medium shots. And when editing footage, the model sometimes paints in a detail from later in the source — in one cut, a shaved line in a haircut showed up a shot early and vanished in the next. Retries did not clear that one for us, so we cut the piece out.

05How to run it

From prompt to clip in three steps

  1. Describe the shot, not the topic

    Subject, what it is doing, the camera, the light. "Single continuous shot" is worth saying out loud. Then add a short "Audio/sound:" line for what should be heard — Wan takes the hint.

  2. Draft at 480p first

    Four seconds at 480p costs $0.20. Get the composition right there, then re-run the winner at 720p or 1080p for the length you actually need.

  3. Pay for the seconds you got

    Your balance is charged when the job starts, and refunded automatically if it fails or the filter rejects it. Top up with crypto or a card; there is no monthly fee waiting for you.

06Prices

What a clip actually costs

Wan 3.0 by resolution and length, next to the models it competes with
Model4 s10 s30 sSound
Wan 3.0 · 480p$0.05/s — drafts and social$0.20$0.50$1.50Included, free
Wan 3.0 · 720p$0.10/s — the usual choice$0.40$1.00$3.00Included, free
Wan 3.0 · 1080p$0.20/s — the sharp one$0.80$2.00$6.00Included, free
Veo 3.1 Fast · 720pGoogle's model, 4–8 s only$0.35Costs extra
Gemini Omni 1.1 Flash · 720pGoogle's model, by the second, 3–10 s a run$0.40$1.00Always on
Gemini Omni 1.1 Flash · 1080pThe same model at 1080p, $0.15 a second$0.60$1.50Always on

Prices are per finished clip and include the audio track. A dash means the model does not deliver that length in a single run: Veo stops at eight seconds, and Omni reaches 40 seconds only by extending a finished clip scene by scene. Deposits of $50 add 5% to your balance and $100 adds 10%; an active promo code adds another 10% on top of the same deposit. The +5% / +10% volume bonus applies to crypto top-ups only.

The full table for every model lives on the pricing page.

07Read next

Guides that go deeper

08FAQ

Questions people actually ask

How long can a Wan 3.0 video be?

Up to 30 seconds in a single generation. The lengths you can pick are 4, 6, 8, 10, 15, 20 and 30 seconds, and the clip comes back as one continuous take — it is not four short clips joined together. Google's Veo and Omni models stop at eight and ten seconds per clip.

How much does a Wan 3.0 clip cost?

You pay by the second: $0.05 at 480p, $0.10 at 720p, $0.20 at 1080p. So four seconds of draft is $0.20, ten seconds at 720p is $1.00, and the maximum — thirty seconds at 1080p — is $6.00. The audio track is included in that price.

Read the full answer
Does Wan 3.0 generate sound?

Yes, and it is on by default. Sound is rendered together with the picture rather than added afterwards, so footsteps land on the footfall. You can switch it off for a silent clip, and that does not change what you pay. Describe what you want to hear in a short "Audio/sound:" line at the end of the prompt. Worth knowing: Alibaba's own launch notes list audio texture as an area still being refined. In our runs the track came back usable as it is, and spoken lines sounded decent, but on a clip that matters it is worth checking the level.

Can I animate my own photo?

Yes. Drop an image into the starting frame slot and the prompt describes what should happen to it. You can add a closing frame to control where the shot lands, and using the same picture in both slots gives you a loop. One catch: frames and reference images cannot be used together.

Wan 3.0 or Veo 3.1 — which should I use?

Wan 3.0 is the newer and more capable of the two: takes of up to 30 seconds against eight, sound included in the price, reference clips on input, and a product that stays itself when you animate it from a first frame. Veo 3.1 is the alternative for what Wan does not cover — 4K output — and for short clips in Google's look. My rule of thumb: start with Wan, and switch to Veo when the brief says 4K.

Can I edit a clip after it is generated?

Yes, with a caveat. Hand Wan a clip of up to 15 seconds — one it made or one you filmed — describe the change, and it returns a new clip. Source and result together have to fit inside 30 seconds, and the source seconds are billed at the same per-second rate as the result. It is not a conversation: every edit is a fresh request with no memory of the last one, and "extend" gives you the next scene rather than a frame-accurate continuation. Gemini Omni Flash is the model that remembers the clip it made.

Can I use Wan 3.0 videos commercially?

Yes. What you generate is yours to use, including in client and commercial work, subject to the content rules in our terms. There is no watermark on the output and no separate licence to buy.

Thirty seconds, one prompt

New accounts start with $0.20 on the balance — enough for a 480p draft to see whether Wan is your model. Clips from $0.20, no subscription, no commitment.