The AI Video Prompt Guide: Camera Motion Without the Melting
AI video fails in predictable ways. Learn the one-move rule, why image-to-video beats text-to-video, and how to write a shot that survives eight full seconds.
Text-to-video models are astonishing for four seconds and unstable for ten. Almost every artefact you have seen — melting hands, morphing faces, backgrounds that breathe — traces back to asking for too much movement at once.
The one-move rule
Describe exactly one camera move per generation. A slow dolly in. A lateral track. A gentle crane up. The moment you ask for a dolly that also pans and then zooms, the model has to reconcile three motion vectors and it will invent geometry to do it.
If you need a complex move, generate it as two shots and cut between them. That is what a real crew would do anyway.
Start from a still
Image-to-video is dramatically more stable than text-to-video. Generate your hero frame first — get the lighting, wardrobe and composition exactly right — then animate it. You keep full control over the look, and the model only has to solve motion.
Write motion in layers
A good video prompt separates four things:
- Camera — the move, its speed, where it starts and ends.
- Subject motion — what the subject actually does, described as a single continuous action.
- Ambient motion — hair, fabric, steam, dust. This is what makes a shot feel alive.
- Grade — contrast, colour, film response.
Keeping them separate stops the model from applying your camera instruction to the subject, which is a common source of drift.
Slow beats fast
Fast motion is where artefacts appear. A constant, slow move gives the model time to keep its geometry consistent across frames. If you need energy, get it from the content of the frame, not the speed of the camera.
Say what you do not want
Video negative prompts are short and specific: jump cuts, flickering, morphing hands, text overlays, sudden zoom, frame stutter. These are the actual failure modes. Listing "ugly, bad quality" does nothing.
Finish with sound
Silent AI video reads as fake even when the picture is flawless. Thirty seconds of room tone, a wind bed or a simple music cue will do more for perceived quality than another five generations.