How do you write AI video prompts that move the way you want?
Describe the motion, not the picture: one scene, one subject, one camera move, sound last. Starting from a still, write only what moves.
A good AI video prompt describes motion, not a picture. Keep to one scene and one subject, give the camera one idea, and put sound at the end. When you start from a still, the still already decides the look, so the prompt only says what moves. Everything about composition, colour and text is in our image prompt guide.
What is the difference between text to video and image to video prompts?
Pick image to video when the look has to stay exact: a product, a character you already drew, a frame you approved. Pick text to video when you are still exploring, and be specific, because every word you leave out is a decision the model makes for you. Models differ in how they use your picture. On some it becomes the first frame of the clip; on others it is a reference the model draws from, so the clip can open on a different framing.
| Text to video | Image to video | |
|---|---|---|
| What you give | A prompt | A picture and a prompt |
| What the prompt describes | Subject, setting, light, action and camera | Only what moves and how the camera behaves |
| Who decides the look | The model, guided by your words | Your picture |
| Best for | Ideas, moods and scenes you have not drawn yet | Products, characters and frames that must stay as they are |
| On Seedance 1.5 Pro | Prompt only | The picture becomes the first frame |
| On the Seedance 2.0 models | Prompt only | The picture is a reference, not the first frame |

A photorealistic vertical wide shot of a small whitewashed lighthouse keeper's cottage standing on the edge of a grassy sea cliff at dusk. Warm amber light glows from its two small windows. Far below, dark blue waves break against the rocks in white foam. The sky fades from deep violet overhead to a thin band of orange at the horizon, with soft scattered clouds. A few gulls hang in the air near the cliff. No people. Calm, quiet, cinematic mood, natural colours.
The waves below roll in and break against the rocks, white foam spreading and pulling back. The window light flickers softly. Clouds drift slowly to the right across the dusk sky. Two gulls glide past the cliff and bank away. Static camera.
A small whitewashed lighthouse keeper's cottage stands on the edge of a grassy sea cliff at dusk. Warm amber light glows from its two small windows. Far below, dark blue waves roll in and break against the rocks, white foam spreading and pulling back. The sky fades from deep violet to a thin orange band at the horizon, and soft clouds drift slowly to the right. Two gulls glide past the cliff and bank away. Wide shot, static camera, photorealistic, calm mood.
How do you describe camera movement and angle?
An angle is where the camera stands; a move is what it does during the clip. You can combine one of each, such as a low side angle with a tracking move. Two moves at once, a pan into a zoom into an orbit, is the most common reason a clip jitters or drifts. If you need a second move, make a second clip and join them in an editor.
| Shot | Angle or move | What it does | Example line |
|---|---|---|---|
| Frontal | Angle | Faces the subject straight on | Camera: frontal shot at eye level. |
| Side | Angle | Shows the subject in profile, good for movement across the frame | Camera: direct side shot as she walks past. |
| Overhead | Angle | Looks straight down, good for tables, food and hands | Camera: overhead shot of the cutting board. |
| Rear | Angle | Follows the subject's view from behind | Camera: rear shot as he walks toward the sea. |
| Close-up | Angle | Fills the frame with a face, hands or a product | Camera: close-up of her hands on the clay. |
| Pan | Move | Turns left or right across the scene | Camera: slow pan from left to right along the stalls. |
| Glide | Move | Slides smoothly sideways or up, without turning | Camera: smooth glide up past the shelves. |
| Tilt | Move | Turns up or down from a fixed spot | Camera: slow tilt up from the entrance to the roof. |
| Zoom | Move | Closes in on or pulls back from one detail | Camera: slow zoom in to a close-up of her face. |
| Orbit | Move | Circles around the subject | Camera: slow orbit around the bottle. |
| Tracking | Move | Travels alongside or behind a moving subject | Camera: tracking shot alongside the cyclist. |
| Static | Move | Does not move, so only the subject does | Camera: static wide shot. |
A busy outdoor street market in the morning: stalls piled with oranges, peppers and fresh bread under striped awnings, shoppers walking between them, sellers handing over paper bags. Soft warm sunlight. The camera pans slowly from left to right along the row of stalls at eye level. Photorealistic.
A tall glass skyscraper on a clear afternoon, its windows reflecting blue sky and passing clouds. The camera starts at street level on the entrance and tilts slowly upward along the glass facade until it reaches the top of the tower against the sky. Photorealistic.
A young runner crossing a finish line on a city road, sweat on her forehead, breathing hard, then breaking into a tired smile. Late afternoon light, blurred crowd behind her. The camera starts on a medium shot and zooms in slowly to a close-up of her face. Photorealistic.
A plain glass perfume bottle with amber liquid standing on a rough grey stone slab, soft studio light, a dark background. Light glints and moves across the glass as the camera orbits slowly and smoothly around the bottle in a half circle. Photorealistic product shot.
A cyclist in a red jersey rides along a winding coastal road above a blue sea on a sunny morning, legs pushing steadily on the pedals. The camera tracks alongside him at the same speed, keeping him in the middle of the frame as the sea and cliffs slide past behind. Photorealistic.
A dancer in a flowing white dress spins on the spot in the middle of an empty wooden studio, the skirt lifting and swirling around her, then settling as she slows. Soft daylight from tall windows. Static camera, locked off, wide shot. Photorealistic.
How do you make motion look real?
Motion looks fake when nothing reacts to it. So after the action, write its result: the syrup pools, spreads to the edges, soaks in and runs down the sides in slow drips. Write in the present tense, with plain verbs and a speed word where it matters. Describe the main person or object once, clearly, and do not hand the action to a second subject halfway through. Small background motion, such as smoke, leaves or passing people, makes a scene feel alive without competing.
A stack of golden pancakes on a white plate on a kitchen table in soft morning light. Thick amber syrup is poured slowly onto the top pancake; it pools in the middle, spreads out to the edges, soaks in, and runs down the sides in slow drips that gather on the plate. Close-up, static camera. Photorealistic.
How long should a video prompt be?
Our full latte art prompt is 88 words and fills a five-second clip with one pour. One Seedance guide puts the budget under 100 words. A clip that tries to fit a walk, a conversation and a sunset will rush all three. For something longer, write the beats in order, each with its own time range, and keep the same description of the person and the place in every beat so they stay the same.
We wrote one idea, a barista pouring latte art, four times. Each version adds one kind of decision: the shot, then the person and the action in small steps, then the place, the camera move and the light. Each version was drawn as its own five-second clip.
| Step | What the step adds | Length |
|---|---|---|
| 1. Bare | only the subject and the action | 5 words |
| 2. Shot | only the shot type | 8 words |
| 3. Detail | who the person is and the action in small physical steps | 61 words |
| 4. Full | the place, one camera move and the light | 88 words |
A barista makes latte art.
Close-up shot of a barista making latte art.
Close-up shot of a young barista with short dark hair and a green apron, focused expression. He tilts a white cup of espresso, lowers a steel milk jug close to the surface, pours a thin stream of steamed milk, wiggles the jug gently from side to side, then lifts it and draws a line through the middle to finish a heart.
Close-up shot of a young barista with short dark hair and a green apron, focused expression, in a small cafe with wooden counters and plants. Warm morning sunlight comes through the window beside him. He tilts a white cup of espresso, lowers a steel milk jug close to the surface, pours a thin stream of steamed milk and wiggles the jug gently from side to side; the white foam spreads across the brown crema and folds into a heart. The camera slowly pushes in toward the cup. Photorealistic.
Should you describe sound in the prompt?
Keep sound after the picture. The visual instructions are the ones you need followed, so they go first, and the sound line is short: one or two named effects and a mood for the music, or none. A spoken line or a script, such as "tap to learn more", tends to come out garbled and steals words from the scene. Record or generate the voice separately and lay it over the clip in an editor.
Which Seedance model should you write for?
The same day Seedance 2.0 Mini rendered a test clip, and Seedance 2.0 Fast with a reference picture timed out after 30 minutes. On Seedance 1.5 Pro the picture became the first frame of the clip; on the 2.0 models it was used as a reference.
About 2 minutes on Seedance 1.5 Pro, 4 on Seedance 2.0 Mini, 6 on Seedance 2.0 Fast and about an hour on Seedance 2.0. Queue times change with load; these are what we saw on one day.
| Model | What your picture does | Queue time we saw | Our run |
|---|---|---|---|
| Seedance 1.5 Pro | Becomes the first frame | About 2 minutes | 8 of 8 clips from our pictures rendered |
| Seedance 2.0 Mini | Used as a reference | About 4 minutes | A test clip rendered |
| Seedance 2.0 Fast | Used as a reference | About 6 minutes | With a reference picture, timed out after 30 minutes |
| Seedance 2.0 | Used as a reference | About an hour | Not part of this run |
Measured in the N33 AI Studio on 9 October 2026.
| Mistake | Fix |
|---|---|
| Talking to the model: "make it faster than last time" | Describe what should be on screen, as if for the first time |
| Several scenes in one clip | One scene per clip, then join the clips in an editor |
| A pan, a zoom and an orbit together | One camera move per clip |
| Describing the still again in an image to video prompt | Write only what moves and what the camera does |
| A vague verb: "the car moves" | Say what moves, where, how fast, and what it causes |
| Too many events for a short clip | One beat per clip, longer ideas beat by beat |
| An ad script or a line of dialogue | Describe the scene; add words and voice afterwards |
| New text asked to appear on screen | Add captions in an editor; if you must, keep to short common words in the prompt's language |
| A reference picture with no job | Say what it is for: the face, the product, the setting |
| The face drifts during the clip | Describe the person once in detail and start from a picture of them |
To try a prompt from this guide, write it on our Seedance page, bring a still to life with image to video, or start from text in the AI video generator. For the picture you start from, see our AI image prompts and the image prompt guide. You can also polish a prompt for free in our chat before you send it.
Questions and answers
How do you keep a face the same across a clip?
Start from a picture of the person rather than from text, describe them once in detail (hair, clothing, one or two features) and ask for the same face throughout. Keep the action small and the camera simple: fast turns and big moves are where faces change. Use a generated person or yourself, never someone who has not agreed to it.
How do you add a seasonal touch to a video?
Add one or two small cues rather than a new scene: snow drifting past a window, string lights in the background, a pumpkin on the counter, a scarf on the person. Put them in the setting sentence so they stay in the background and do not take over the action.
Can AI video put text on screen?
Sometimes, but it is the least reliable part of a clip. If you need it, keep it to a few common words in the same language as the prompt, with no symbols. For anything that has to be spelled right, such as a name or an offer, add the text in an editor afterwards.
How long can an AI video be?
Each generation makes one short clip, and every clip in this guide is five seconds. Longer videos are built from several clips, each with its own prompt, joined in an editor. Keep each prompt focused, since one Seedance guide finds instructions past about 100 words are followed less.