Five seconds is one beat, not five
Most bad generated acting is three performances stacked into one take. Writing action is mostly subtraction.

The most common note on a generated performance is that the character is doing too much. Not badly. Too much. Five seconds contains a look down, a sigh, a head shake, a hand gesture and a swallow, because the request asked for emotion and the model delivered all of it at once.
Real acting in a five second shot is usually one thing, occasionally two. This is a method for getting there, with the takes to compare against.
Write beats, not feelings
A feeling is not directable. Ask for conflicted and you are asking the model to invent a physical vocabulary for an abstraction, which it will do enthusiastically and generically. A beat is a physical event with a start and an end: he looks down at his hands, he breathes out, he lifts his eyes.
Three beats in five seconds is the ceiling. Here is exactly that, with an explicit instruction that there be no gesture and no head shake.
Know what too much looks like
Worth generating the failure deliberately once, so you can recognise it in work you did not intend to make. This is the same man, same room, same camera, asked for a broadly emotional performance instead of specific beats.
Nothing here is a rendering fault. It is a direction fault, and it is what models default to because their priors about performance are broad and theatrical.
Most bad generated acting is not a failure of expression. It is three performances stacked into one take.
Say what you do not want
Negative direction is the highest-leverage clause in the whole request, and almost nobody writes it. No gesture. No head shake. Small movements only. Does not speak. Each of those removes an entire family of default behaviour, and what is left is much closer to how people actually sit in a room.
Compare the two clips above and this is the only meaningful difference in the instruction. Not the framing, not the light, not the length.
Cast the take, do not fix the prompt
Performance has variance in a way that framing does not. The difference between the first take and the best of four is usually larger than any prompt edit you could make, so generate the same beats several times and choose, exactly as you would on a set.
Neither is wrong. One will cut better than the other, and which one depends entirely on the shot before it. Watch them without the sound of your own intentions in your head; the take that acts best is rarely the one that follows the description most literally.
Direct the eyes
Where a character looks does more narrative work in five seconds than anything their face is doing. Looking at an empty chair implies an absence. Looking past the camera implies someone standing there. Looking down and staying down implies a decision already made.


Specify the moment a look breaks, too. A held gaze that drops at the end of a clip means something different from one that lifts, and if you do not choose, the model chooses for you.
Lock the camera and end on a hold
A moving camera competes with a face. If the interesting thing is what a character is doing, the camera should stop having opinions: locked off, static, no push, no zoom. This is the opposite of the reflex, because movement reads as production value, and the push then pulls attention off the only thing in frame worth watching.
Then let the clip end on a hold rather than an action. Ending on movement cuts mid-gesture, makes the loop point obvious, and reads as a fragment. A final hold leaves the shot open, which is what makes five seconds feel like part of a scene.
Breath is the timing device that makes stillness work. Breathes out slowly is an instruction with a duration attached, roughly a second and a half. A character asked to hold completely still reads as frozen; a character breathing is still, but alive, and that gap is most of what separates a usable take from an uncanny one.
