How do you write AI video prompts that move the way you want?

Published October 9, 2026 · Updated October 9, 2026

Describe the motion, not the picture: one scene, one subject, one camera move, sound last. Starting from a still, write only what moves.

A good AI video prompt describes motion, not a picture. Keep to one scene and one subject, give the camera one idea, and put sound at the end. When you start from a still, the still already decides the look, so the prompt only says what moves. Everything about composition, colour and text is in our image prompt guide.

What is the difference between text to video and image to video prompts?

In text to video the prompt builds everything: subject, setting, light and action. In image to video your picture already fixes the look, so the prompt describes only the motion and the camera. Repeating the picture in words adds nothing and can pull the clip away from it.

Pick image to video when the look has to stay exact: a product, a character you already drew, a frame you approved. Pick text to video when you are still exploring, and be specific, because every word you leave out is a decision the model makes for you. Models differ in how they use your picture. On some it becomes the first frame of the clip; on others it is a reference the model draws from, so the clip can open on a different framing.

Text to video and image to video, side by side
Text to videoImage to video
What you giveA promptA picture and a prompt
What the prompt describesSubject, setting, light, action and cameraOnly what moves and how the camera behaves
Who decides the lookThe model, guided by your wordsYour picture
Best forIdeas, moods and scenes you have not drawn yetProducts, characters and frames that must stay as they are
On Seedance 1.5 ProPrompt onlyThe picture becomes the first frame
On the Seedance 2.0 modelsPrompt onlyThe picture is a reference, not the first frame
A whitewashed lighthouse keeper's cottage on a grassy sea cliff at dusk
The still we started from. Light, colour and framing are already decided here, so the video prompt does not repeat them.
PromptNano Banana
A photorealistic vertical wide shot of a small whitewashed lighthouse keeper's cottage standing on the edge of a grassy sea cliff at dusk. Warm amber light glows from its two small windows. Far below, dark blue waves break against the rocks in white foam. The sky fades from deep violet overhead to a thin band of orange at the horizon, with soft scattered clouds. A few gulls hang in the air near the cliff. No people. Calm, quiet, cinematic mood, natural colours.
5 s
Image to video from that still: the prompt names only the waves, the window light, the clouds, the gulls and a still camera.
PromptSeedance 1.5 Pro
The waves below roll in and break against the rocks, white foam spreading and pulling back. The window light flickers softly. Clouds drift slowly to the right across the dusk sky. Two gulls glide past the cliff and bank away. Static camera.
5 s
Text to video of the same idea: with no picture, the prompt has to carry the place, the light and the camera too.
PromptSeedance 1.5 Pro
A small whitewashed lighthouse keeper's cottage stands on the edge of a grassy sea cliff at dusk. Warm amber light glows from its two small windows. Far below, dark blue waves roll in and break against the rocks, white foam spreading and pulling back. The sky fades from deep violet to a thin orange band at the horizon, and soft clouds drift slowly to the right. Two gulls glide past the cliff and bank away. Wide shot, static camera, photorealistic, calm mood.

How do you describe camera movement and angle?

Give the camera its own short sentence, after the action: the angle it sees from and the one move it makes, with a speed word such as slow or steady. Use plain film terms like pan, tilt, zoom, orbit, tracking or static, and one move per clip.

An angle is where the camera stands; a move is what it does during the clip. You can combine one of each, such as a low side angle with a tracking move. Two moves at once, a pan into a zoom into an orbit, is the most common reason a clip jitters or drifts. If you need a second move, make a second clip and join them in an editor.

Camera angles and moves, with one line to copy for each
ShotAngle or moveWhat it doesExample line
FrontalAngleFaces the subject straight onCamera: frontal shot at eye level.
SideAngleShows the subject in profile, good for movement across the frameCamera: direct side shot as she walks past.
OverheadAngleLooks straight down, good for tables, food and handsCamera: overhead shot of the cutting board.
RearAngleFollows the subject's view from behindCamera: rear shot as he walks toward the sea.
Close-upAngleFills the frame with a face, hands or a productCamera: close-up of her hands on the clay.
PanMoveTurns left or right across the sceneCamera: slow pan from left to right along the stalls.
GlideMoveSlides smoothly sideways or up, without turningCamera: smooth glide up past the shelves.
TiltMoveTurns up or down from a fixed spotCamera: slow tilt up from the entrance to the roof.
ZoomMoveCloses in on or pulls back from one detailCamera: slow zoom in to a close-up of her face.
OrbitMoveCircles around the subjectCamera: slow orbit around the bottle.
TrackingMoveTravels alongside or behind a moving subjectCamera: tracking shot alongside the cyclist.
StaticMoveDoes not move, so only the subject doesCamera: static wide shot.
5 s
Pan: the camera turns along the stalls while the shoppers keep walking.
PromptSeedance 1.5 Pro
A busy outdoor street market in the morning: stalls piled with oranges, peppers and fresh bread under striped awnings, shoppers walking between them, sellers handing over paper bags. Soft warm sunlight. The camera pans slowly from left to right along the row of stalls at eye level. Photorealistic.
5 s
Tilt: from the entrance up to the top of the tower, the only movement in the shot.
PromptSeedance 1.5 Pro
A tall glass skyscraper on a clear afternoon, its windows reflecting blue sky and passing clouds. The camera starts at street level on the entrance and tilts slowly upward along the glass facade until it reaches the top of the tower against the sky. Photorealistic.
5 s
Zoom: from a medium shot to her face, so the clip ends on the detail that matters.
PromptSeedance 1.5 Pro
A young runner crossing a finish line on a city road, sweat on her forehead, breathing hard, then breaking into a tired smile. Late afternoon light, blurred crowd behind her. The camera starts on a medium shot and zooms in slowly to a close-up of her face. Photorealistic.
5 s
Orbit: the camera circles the bottle and the light slides across the glass.
PromptSeedance 1.5 Pro
A plain glass perfume bottle with amber liquid standing on a rough grey stone slab, soft studio light, a dark background. Light glints and moves across the glass as the camera orbits slowly and smoothly around the bottle in a half circle. Photorealistic product shot.
5 s
Tracking: the camera keeps pace with the cyclist, so the road and the sea move behind him.
PromptSeedance 1.5 Pro
A cyclist in a red jersey rides along a winding coastal road above a blue sea on a sunny morning, legs pushing steadily on the pedals. The camera tracks alongside him at the same speed, keeping him in the middle of the frame as the sea and cliffs slide past behind. Photorealistic.
5 s
Static: the camera stays still and all the motion belongs to the dancer and her dress.
PromptSeedance 1.5 Pro
A dancer in a flowing white dress spins on the spot in the middle of an empty wooden studio, the skirt lifting and swirling around her, then settling as she slows. Soft daylight from tall windows. Static camera, locked off, wide shot. Photorealistic.

How do you make motion look real?

Break the action into small physical steps, write what each step causes, and keep one subject in charge. "Pours coffee" is one vague verb; tilting the cup, a thin stream and foam spreading into a heart are three things the model can actually draw.

Motion looks fake when nothing reacts to it. So after the action, write its result: the syrup pools, spreads to the edges, soaks in and runs down the sides in slow drips. Write in the present tense, with plain verbs and a speed word where it matters. Describe the main person or object once, clearly, and do not hand the action to a second subject halfway through. Small background motion, such as smoke, leaves or passing people, makes a scene feel alive without competing.

5 s
Cause and effect: the prompt says what the syrup does after it lands, not only that it is poured.
PromptSeedance 1.5 Pro
A stack of golden pancakes on a white plate on a kitchen table in soft morning light. Thick amber syrup is poured slowly onto the top pancake; it pools in the middle, spreads out to the edges, soaks in, and runs down the sides in slow drips that gather on the plate. Close-up, static camera. Photorealistic.

How long should a video prompt be?

One beat per short clip: a subject, one action with its result, a camera line and maybe a sound line. Put the subject and the action first, since instructions near the end of a long prompt are followed least. For a longer clip, write it beat by beat.

Our full latte art prompt is 88 words and fills a five-second clip with one pour. One Seedance guide puts the budget under 100 words. A clip that tries to fit a walk, a conversation and a sunset will rush all three. For something longer, write the beats in order, each with its own time range, and keep the same description of the person and the place in every beat so they stay the same.

What we measured ourselves
2026-10-09 → 2026-10-09
5 to 88 wordswords in the latte art prompt, from the one-line first try to the full version

We wrote one idea, a barista pouring latte art, four times. Each version adds one kind of decision: the shot, then the person and the action in small steps, then the place, the camera move and the light. Each version was drawn as its own five-second clip.

Source: N33 engines, staff run
Four versions of one video prompt, and how long each was
StepWhat the step addsLength
1. Bareonly the subject and the action5 words
2. Shotonly the shot type8 words
3. Detailwho the person is and the action in small physical steps61 words
4. Fullthe place, one camera move and the light88 words
5 s
Step one: just the subject and the action. The model picks the shot, the person, the place and the camera on its own.
PromptSeedance 1.5 Pro
A barista makes latte art.
5 s
Step two: only the shot type, a close-up. Everything else is still the model's choice.
PromptSeedance 1.5 Pro
Close-up shot of a barista making latte art.
5 s
Step three: the person and the pour broken into small physical moments, which is what makes the hands move believably.
PromptSeedance 1.5 Pro
Close-up shot of a young barista with short dark hair and a green apron, focused expression. He tilts a white cup of espresso, lowers a steel milk jug close to the surface, pours a thin stream of steamed milk, wiggles the jug gently from side to side, then lifts it and draws a line through the middle to finish a heart.
5 s
Step four: the same action with a place, a single slow push in and the light. Nothing in it fights the action for attention.
PromptSeedance 1.5 Pro
Close-up shot of a young barista with short dark hair and a green apron, focused expression, in a small cafe with wooden counters and plants. Warm morning sunlight comes through the window beside him. He tilts a white cup of espresso, lowers a steel milk jug close to the surface, pours a thin stream of steamed milk and wiggles the jug gently from side to side; the white foam spreads across the brown crema and folds into a heart. The camera slowly pushes in toward the cup. Photorealistic.

Should you describe sound in the prompt?

On models that render sound with the picture, yes, as its own last line: name the sounds close to their source, such as milk hissing or waves on rocks, or write that there is no music. Leave dialogue and voiceover out and add them afterwards.

Keep sound after the picture. The visual instructions are the ones you need followed, so they go first, and the sound line is short: one or two named effects and a mood for the music, or none. A spoken line or a script, such as "tap to learn more", tends to come out garbled and steals words from the scene. Record or generate the voice separately and lay it over the clip in an editor.

Which Seedance model should you write for?

The prompt grammar is the same on every Seedance model. What changes is how your picture is used and how long you wait: Seedance 1.5 Pro starts the clip from your picture, the 2.0 models treat it as a reference.
What we measured ourselves
2026-10-09 → 2026-10-09
8 of 8clips rendered by Seedance 1.5 Pro from our own pictures

The same day Seedance 2.0 Mini rendered a test clip, and Seedance 2.0 Fast with a reference picture timed out after 30 minutes. On Seedance 1.5 Pro the picture became the first frame of the clip; on the 2.0 models it was used as a reference.

Source: N33 Studio, staff run
What we measured ourselves
2026-10-09 → 2026-10-09
about 2 minutes to about an hourtime a clip waited in our Studio queue, by model

About 2 minutes on Seedance 1.5 Pro, 4 on Seedance 2.0 Mini, 6 on Seedance 2.0 Fast and about an hour on Seedance 2.0. Queue times change with load; these are what we saw on one day.

Source: N33 Studio, staff run
The Seedance models in our Studio, as we measured them
ModelWhat your picture doesQueue time we sawOur run
Seedance 1.5 ProBecomes the first frameAbout 2 minutes8 of 8 clips from our pictures rendered
Seedance 2.0 MiniUsed as a referenceAbout 4 minutesA test clip rendered
Seedance 2.0 FastUsed as a referenceAbout 6 minutesWith a reference picture, timed out after 30 minutes
Seedance 2.0Used as a referenceAbout an hourNot part of this run

Measured in the N33 AI Studio on 9 October 2026.

Common video prompt mistakes and how to fix them
MistakeFix
Talking to the model: "make it faster than last time"Describe what should be on screen, as if for the first time
Several scenes in one clipOne scene per clip, then join the clips in an editor
A pan, a zoom and an orbit togetherOne camera move per clip
Describing the still again in an image to video promptWrite only what moves and what the camera does
A vague verb: "the car moves"Say what moves, where, how fast, and what it causes
Too many events for a short clipOne beat per clip, longer ideas beat by beat
An ad script or a line of dialogueDescribe the scene; add words and voice afterwards
New text asked to appear on screenAdd captions in an editor; if you must, keep to short common words in the prompt's language
A reference picture with no jobSay what it is for: the face, the product, the setting
The face drifts during the clipDescribe the person once in detail and start from a picture of them

To try a prompt from this guide, write it on our Seedance page, bring a still to life with image to video, or start from text in the AI video generator. For the picture you start from, see our AI image prompts and the image prompt guide. You can also polish a prompt for free in our chat before you send it.

Questions and answers

How do you keep a face the same across a clip?

Start from a picture of the person rather than from text, describe them once in detail (hair, clothing, one or two features) and ask for the same face throughout. Keep the action small and the camera simple: fast turns and big moves are where faces change. Use a generated person or yourself, never someone who has not agreed to it.

How do you add a seasonal touch to a video?

Add one or two small cues rather than a new scene: snow drifting past a window, string lights in the background, a pumpkin on the counter, a scarf on the person. Put them in the setting sentence so they stay in the background and do not take over the action.

Can AI video put text on screen?

Sometimes, but it is the least reliable part of a clip. If you need it, keep it to a few common words in the same language as the prompt, with no symbols. For anything that has to be spelled right, such as a name or an offer, add the text in an editor afterwards.

How long can an AI video be?

Each generation makes one short clip, and every clip in this guide is five seconds. Longer videos are built from several clips, each with its own prompt, joined in an editor. Keep each prompt focused, since one Seedance guide finds instructions past about 100 words are followed less.