How to Use Reference to Video in Wan 3.0
Reference to video in Wan 3.0 means starting from a clip instead of text or a still image: you upload a reference video, then describe what the new clip should inherit (subject action, camera movement, style) and what it should change (scene, wardrobe, lighting). On wan3.dev this sits under the Video to video start option, alongside Text to video, Image to video, and Video editing. It fits you best when you already have footage whose motion or look you want to reuse but whose content you want to replace. If you only have a written idea or a single photo, text-to-video or image-to-video is the cheaper, simpler path.
What the reference actually controls
Wan 3.0's generator lists three ways to start — Text / Image / Video — and the site describes the workflow showcase as covering "subject action, camera direction, image-led motion, and visual transformation." In practice, a reference clip is your strongest lever on:
- Motion — how the subject moves and how fast the action reads
- Camera language — pans, pushes, handheld feel, framing changes
- Style and grade — color, contrast, texture, overall look
- Transformation direction — what the clip turns into over its duration
What it does not reliably lock down is identity and exact detail. A clean reference gives the model a clear motion signal; a busy or low-quality one gives it noise, and the output drifts.
Preparing the reference clip
The site's stated output ceiling is up to 1080p and up to 30 seconds, so treat that as the target range for what you feed in.
| Reference quality | Effect on the result |
|---|---|
| Single clear subject, simple background | Motion and camera transfer most faithfully |
| Multiple subjects or fast cuts | Model averages the motion; output looks muddled |
| Heavy motion blur or low resolution | Weak motion signal, softer output |
| Long clip with several distinct shots | Only part of the motion is reflected; trim to one shot |
Practical prep: trim to one continuous shot, keep the subject large enough in frame to read, and avoid clips where the camera and the subject both move hard at once. If your reference is longer than the clip you want, cut it to roughly the length you plan to generate.
Writing the prompt: separate "keep" from "change"
The prompt field on the generator accepts up to 2000 characters, with a tip to describe subject, style, lighting, mood, and composition. For reference-to-video, split your prompt into two explicit halves:
- Inherit — name the motion and camera you want carried over: "keep the slow push-in and the walking pace of the reference."
- Replace — name everything else: new setting, wardrobe, time of day, color palette.
Example prompt structure:
Keep the reference's forward tracking shot and the subject's steady walking rhythm. Change the setting to a rain-soaked night market, swap the outfit to a red jacket, shift the grade to cool blue with warm practical lights, keep the framing centered.
Vague prompts like "make it like the reference but different" give the model nothing to hold onto; the output tends to copy the reference too literally or ignore it entirely.
Settings and cost before you generate
The generator exposes model selection (Wan 3.0), resolution ratio, duration, a generate-audio toggle, and a thinking mode. The listed range is 2–30 seconds at 480p / 720p / 1080p, costing 98–5880 credits depending on those choices — so both length and resolution move the price, and the low end of that range is a short low-resolution clip while the high end is a long 1080p one. The interface shows a live credit figure next to the Generate button (the page displays an example of 455 credits for one configuration), so set duration and resolution first and read the number before committing.
For a first reference-to-video attempt, generate short and at a lower resolution to confirm the motion transfers the way you expect, then re-run at 1080p once the prompt is right. That keeps failed attempts cheap. Note that pricing and any discounts are shown on the site's own Pricing page — check there rather than assuming a rate.
When the output drifts from your reference
Most failures fall into three buckets:
- Motion ignored — the clip looks static or generic. Your reference likely had weak or ambiguous motion. Swap in a clip with one obvious, sustained movement.
- Reference copied too literally — you get the same scene with minor changes. Your "replace" instructions were too soft. State the new scene, wardrobe, and lighting as explicit changes.
- Muddy result — usually a busy reference or too many simultaneous changes. Simplify to one inherited motion plus one or two replacements, then build up.
Change one variable per retry: fix the reference first, then the prompt, then the settings. Changing all three at once tells you nothing about which one caused the improvement.
A quick decision check
Use reference-to-video when you have a clip whose movement or look is the point and you want new content in that mold. Use image-to-video when a single frame defines the shot and you only need it to move. Use text-to-video when you have no source material and want the model to invent the motion. And if your goal is to rework an existing clip rather than generate a new one from it, the Video editing option is the closer match.