Cinematic Control

Camera Movement and Cinematic Control in Prompts

Share This Spread Love
Rate this post

Most of the difference between generated video that looks amateur and generated video that looks directed comes down to the camera. The subject can be well described, the lighting specified, the action clear — and the clip still reads as a screensaver, because the camera is doing nothing in particular or doing several things at once with no apparent reason.

Camera language is also the part of prompting where precision pays off most steeply, and Seedance 2.0 rewards it more than any model I have used. Vague descriptions of subjects produce reasonable results, because the model has a large space of plausible subjects to work with. Vague descriptions of camera behaviour produce arbitrary results, because there is no default that is right for most shots.

Technical terms outperform descriptions

The first thing worth internalising is that the vocabulary of cinematography is not decoration. Those words exist because they name specific, unambiguous operations, and Seedance 2.0 responds to them far more reliably than to descriptions of the same thing in ordinary language.

“The camera goes around the object” is ambiguous — around at what speed, in which direction, at what distance, staying level or arcing? “Slow orbit, counter-clockwise, eye level, holding distance” is not ambiguous, and produces what it says. The same goes across the board: “moves closer” is weaker than “push in”, “the view gets wider” is weaker than “pull back”, “looking down at it” is weaker than “high angle” or “overhead” depending on which is meant.

Shot size is the term most often left out and the one that changes Seedance 2.0 results most. Extreme close-up, close-up, medium shot, medium wide, wide, establishing — naming one of these fixes the framing decision that the model would otherwise make arbitrarily. A prompt that specifies nothing about shot size gets whatever the model considers a natural framing for the described content, which varies between generations and makes a sequence look inconsistent even when everything else is controlled.

Speed is a separate instruction

A movement without a speed attached is half specified, and this took me longer to learn than it should have.

“Push in” and “slow push in” produce noticeably different clips. So does “rapid push in”, which reads as urgency or alarm rather than the contemplative arrival that a slow one produces. The move is the same operation; the speed is what gives it meaning. In practice I now attach a speed qualifier to every Seedance 2.0 camera instruction as a matter of habit — slow, steady, gradual, rapid, sudden — because leaving it out hands over a decision that carries most of the emotional weight of the shot.

Speed also interacts with duration in a way worth planning for. A slow orbit that would take twelve seconds to feel complete will look truncated if it is generated at four. Matching the ambition of the move to the length of the clip prevents the most common form of unsatisfying output: a movement that stops before it has arrived anywhere.

Anyone rebuilding a prompt library assembled before this year should note the platform address changed — Seedance2.ai moved to Seevio.ai, and saved work and account history carried over intact.

One move per shot

The instinct when learning this vocabulary is to use all of it. A prompt requesting a push in, followed by a pan, with a rack focus at the end, will generate something — but what it generates is usually muddled motion rather than a sequence of distinct moves, because a few seconds is not enough time to execute three operations legibly.

Real cinematography follows the same rule for the same reason. Shots that combine moves exist, but they combine two at most and they are motivated by something happening in frame. A camera that pans while pushing in is following a subject who is moving diagonally; it is not doing two things for the sake of it.

So: one move per shot, and if a sequence needs more than one, that is a sequence of shots rather than a single busier shot. Seedance 2.0 can produce multiple shots with cuts inside a single generation, which is a better way to get a varied camera than by loading one continuous take with instructions.

Motivated movement reads better than arbitrary movement

The clips that look directed have a reason for the camera to be doing what it is doing, and stating that reason in the prompt improves results measurably.

“Tracking shot following the subject as she walks left” performs better in Seedance 2.0 than “tracking shot, camera moves left” even though they describe similar operations, because the first tells the model what the movement is anchored to. Camera motion that is tied to something in frame — following a subject, revealing something, arriving at a detail — is generated more coherently than camera motion described as an abstract operation.

This is the single adjustment that most improved my output. Not more precise camera terminology, but attaching the camera to a reason.

Lens language

Beyond movement, the vocabulary of lenses gives control over how a shot feels that shot size alone does not reach.

Depth of field is the most useful, and Seedance 2.0 handles it consistently. “Shallow depth of field” separates a subject from its background and reads as photographic and intentional; “deep focus” holds everything sharp and reads as documentary or observational. Specifying one of these is nearly always better than specifying neither.

Focal length terms work, with caveats. “Wide angle” reliably produces the expansive, slightly distorted perspective it should, and “telephoto” produces the compressed, flattened one. Precise numerical focal lengths are less consistent — I get better results describing the effect than naming a millimetre value.

Rack focus, where focus shifts from one plane to another during the shot, works when the prompt makes clear what starts sharp and what ends sharp. Left vague it usually does not happen at all.

Light and grade

Lighting has been covered at length elsewhere, but two points belong in a discussion of cinematic control specifically.

Contrast ratio carries more of the “cinematic” quality than any camera move, in Seedance 2.0 as anywhere else. High contrast with deep shadows reads as dramatic and film-like; even, flat lighting reads as video regardless of what the camera is doing. Prompts that specify a single directional key with falloff into shadow produce more cinematic results than prompts that specify a move but leave light open.

Colour treatment is worth stating explicitly. Warm and cool as broad directions, desaturated or saturated, high contrast or gentle — these register reliably. Naming specific film stocks, cameras, or directors is less reliable and, for commercial work, carries considerations beyond whether it happens to work.

What remains unreliable

Some things I have stopped attempting.

Very long or complex single takes with multiple beats do not hold together within the durations available; they are better broken into shots.

Precise geometric camera paths — a specific arc through a specific space, ending at a specific mark — are approximated rather than executed. Where exact camera geometry matters, this is not the tool.

Handheld as a specific texture is inconsistent. Sometimes it produces convincing organic movement, sometimes something closer to a wobble, and I have not found phrasing that makes it dependable.

Camera moves through complex environments with a lot of occlusion — passing behind objects, moving through doorways — tend to produce spatial inconsistency at the transition points.

The practical shape of it

What I write now for a controlled shot has a consistent form: subject and action first, then one camera instruction with a speed and an anchor, then shot size, then depth of field, then light with direction and quality, then style register. Six clauses, two or three sentences, nothing left implicit that matters.

The reason this works is not that the model needs a formula. It is that every element left unspecified is a decision the model makes without input, and camera decisions made without input are the ones that most visibly separate a directed clip from a generated one. Seedance 2.0 follows this vocabulary closely enough that the constraint on output is now mostly the specificity of the instruction rather than the capability of the model — which is a good problem to have, and a learnable one.