How to Write Wan 3.0 Video Prompts That Get the Clip You Want
Wan 3.0's prompt box accepts up to 2,000 characters and its own tip tells you to describe the subject, style, lighting, mood, and composition. That is the whole method: name what is on screen, say what it does, say how the camera behaves, then set the light and the look. Prompts written this way work for all three starting points — text, a photo, or a reference clip — and you can adjust one element at a time when a result misses instead of rewriting from scratch.
The prompt pattern
Build every prompt in this order. Later elements matter less if the earlier ones are vague, so fix the top of the list first.
- Subject — who or what is on screen, with enough detail to be specific ("a woman in a yellow raincoat," not "a person").
- Action — the single main thing that happens during the clip.
- Camera — how the shot moves or holds (static, slow push in, tracking alongside).
- Lighting — time of day and light quality (overcast morning, warm lamp light, hard midday sun).
- Mood and style — the feeling and the visual treatment (documentary, soft and quiet, high contrast).
- Composition — framing and where the subject sits in frame (wide shot, centered, subject on the left third).
A working example:
A woman in a yellow raincoat walks toward the camera along a wet pier, slow push in, overcast morning light, calm and slightly melancholic, wide shot with the subject centered, muted color palette.
Compare that to "a woman walking on a pier, cinematic, 4K, beautiful." The second prompt gives the model nothing to act on — no direction of movement, no camera instruction, no light. Vague quality adjectives are the most common reason a clip comes back looking like a generic stock shot.
Adapting the pattern to each input type
The same six elements apply, but what you should spend characters on changes with how you start.
| Start | What the model already has | What your prompt should carry most |
|---|---|---|
| Text to video | Nothing | All six elements — this is the only case where you describe the subject from zero |
| Image to video | Subject, framing, and look from the photo | Action, camera direction, and how much the scene should change |
| Video to video | Existing motion and structure | What to keep, what to transform, and the target style |
For image-to-video, don't re-describe the photo. If the image already shows a person at a desk, write the motion: "she turns her head toward the window, camera slowly drifts right, late afternoon light." Describing the subject again wastes characters and can push the model away from the source image.
For video-to-video, be explicit about the transformation. "Keep the camera movement and timing; change the setting to a snowy street at night, cool blue tones" tells the model what is fixed and what is not.
How much to write
The box holds 2,000 characters, but length is not the goal — coverage is. A prompt that names all six elements usually lands between 40 and 90 words. Below that, you are probably missing camera or lighting. Far above it, you risk contradicting yourself, which is worse than saying less.
Two habits that keep prompts tight:
- One main action per clip. If you want a character to walk in, sit down, and open a laptop, that is three beats competing for a clip that may run as short as a few seconds. Pick the beat you actually need.
- Cut adjectives that don't change pixels. "Stunning," "amazing," and "masterpiece" don't tell the model what to render. "Backlit," "shallow depth of field," and "handheld" do.
Matching prompt detail to your settings
Your prompt and your settings have to agree, or the result will feel wrong even when the model did what you asked.
- Duration. A prompt describing a slow, multi-stage action needs room to play out. On a short clip, describe a single moment instead of a sequence.
- Resolution. Fine detail you ask for — text on a sign, small objects, intricate patterns — is more likely to survive at higher resolution. At lower settings, favor larger shapes and simpler compositions.
- Audio. If you enable audio, say what should be heard. Ambient sound, dialogue, or music are different requests, and silence is a valid one to state.
- Aspect ratio. Write the composition to match the frame. A "wide establishing shot" fights a vertical frame; "centered subject with headroom" suits it.
When the result is wrong, change one thing
Rewriting the whole prompt after a bad clip makes it impossible to know what fixed it. Instead, diagnose by category:
- Wrong subject or wrong action → the top of your prompt is too vague. Add one concrete detail.
- Right subject, wrong feeling → lighting and mood are underspecified. Name the light source and time of day.
- Static or awkward motion → you described a scene but not a camera. Add an explicit camera instruction.
- Composition is off → state the framing and where the subject sits in frame.
- It ignored part of the prompt → you likely have two competing actions. Cut to one.
Change one element, regenerate, and compare. Two or three passes usually get you to a prompt worth saving and reusing with a different subject.
A reusable template
Fill this in and you have a complete prompt:
[Subject with one distinguishing detail] [does one action], [camera instruction], [lighting and time of day], [mood], [framing and subject placement], [style or color treatment].
Keep your best prompts in a note. Because the pattern is modular, you can swap the subject and keep the camera, lighting, and style — which is the fastest way to get a consistent look across several clips.