Drop or upload up to 10 images
JPG, PNG, WEBP up to 10MB
Drop or upload up to 5 videos
MP4, MOV up to 50M
Drop or upload up to 5 audios
MP3, WAV up to 15M
0/10000
Audio
Thinking
Sample Video
New - Wan 3.0

Wan 3.0 AI Video Generator

Turn prompts, stills, and mixed references into 2–30 second clips at 480p–1080p, with optional audio and thinking mode.

Flagship capability

A full 30-second take, with thinking when the shot is hard

Wan 3.0 renders one continuous clip from 2 to 30 seconds — long enough for a setup and a payoff, short enough to test a move at 2 seconds. Thinking spends extra interpretation on stacked camera, blocking, and timing; it does not change length, resolution, or credit cost.

  • Native 2–30s in one pass, in 1-second steps you choose
  • Thinking mode for complex motion and composition
  • Up to 10 images, 5 clips, and 5 audio tracks as reference
Per-second control
Duration
2–30s

Whole-second clips. You set the length; the UI does not auto-pad or shorten.

No 4K
Resolution
480p–1080p

480p, 720p, or 1080p. Image-to-video follows the first frame’s aspect ratio.

Complex shots
Thinking
Thinking mode

Reads stacked camera, blocking, and timing in the prompt before it renders the take.

Mixed media
References
10 / 5 / 5

Images, video, and audio. Video and audio each total 15 seconds or less.

Three shots this model is actually for

  1. Shot 01

    Turn a prompt into a 2–30s video

    Use this when the scene exists in a prompt, not a still — a camera move that has to land, a beat that needs a middle, not a five-second loop you will stitch later.

    Length
    2–30 seconds in 1-second steps. Use the short end to test a move; use the long end when the story needs a setup and a payoff.
    Thinking
    Turn it on when the prompt stacks camera, blocking, and timing. Leave it off for a simple push-in or pan.
    Output
    Draft at 480p (2 credits/s). The same prompt at 1080p is 8 credits/s and slower to return.

    Thinking does not pick the duration for you and does not add credits. You still set the length. It only spends more interpretation on the prompt you wrote.

  2. Shot 02

    First frame to last frame

    Use this when you already know the start and end composition — a product that should finish on a hero angle, a character who walks from a doorway to a window.

    Length
    Pick the time the move actually needs. Two stills do not justify a 30-second pad.
    Thinking
    On if the interpolation is a complex camera or performance; off if you only need a gentle drift between two similar frames.
    Output
    Aspect ratio is Auto: the clip follows the first frame. You still choose 480p, 720p, or 1080p.

    One image is a start frame. A second image is the last frame the model should land on, not a style board. If you have a character sheet and no destination frame, reference-to-video is the better mode.

  3. Shot 03

    Identity, motion, and sound from files

    Use this when the face, the performance, or the voice should come from files you already have, not from adjectives in the prompt.

    Length
    Output can still be 2–30 seconds. That is independent of the 15-second cap on reference video and the 15-second cap on reference audio.
    Thinking
    On when several references have to stay consistent in one take — a face, a wardrobe, and a motion clip described together.
    Output
    Same three tiers. 1080p plus a long take is the slow, expensive combination; prove the lock at 480p first.

    SparkVid will reject an empty reference set — you need at least one image, clip, or audio file. Put identity in images, performance in video, voice or bed in audio, and say how they relate in the prompt. PDFs and webpages are not accepted here.

Pick Wan 3.0 or switch models

All three run from the same prompt box and credit balance. Choose per shot, not per account.

Reach for this when

Wan 3.0

You want the cheapest 1080p 30-second take on SparkVid, a 2-second minimum for tests, and a thinking toggle for stacked camera and blocking. Audio is optional.

2–30s · 480p / 720p / 1080p · 2 / 4 / 8 credits per second · 10 images, 5 videos, 5 audio tracks (15s media cap each).

Reach for this when

Seedance 2.5

The shot is built from a pile of references, not a short prompt. On SparkVid that budget is 30 images, 10 videos, and 10 audio tracks, with 30 seconds of reference media — double Wan’s media cap — at a higher per-second rate.

4–30s · 480p / 720p / 1080p · 6 / 12 / 24 credits per second when no reference video is attached.

Reach for this when

MiniMax H3

Delivery needs native 2K (2560×1440) and stereo audio in one pass — product type, packaging, or a voiced spot that has to survive a close inspect. Clips stop at 15 seconds.

5–15s · 480p / 768p / 2K · native stereo · up to 9 images, 3 clips, and 3 audio tracks.

Wan 3.0 questions people actually ask

How this generator behaves — length, cost, references, and when to switch models.

Run a Wan 3.0 take in the generator above

Set a length, a resolution, and whether thinking is worth the wait. Credits are shown before you generate.

See pricing