Image-to-Video AI Prompts: How to Animate a Still Without Wrecking It

July 4, 2026 By Bank K.

You found the perfect still. The composition is locked, the lighting is right, the subject looks exactly how you imagined. Then you feed it to an image-to-video model, and the first frame is great while everything after it slowly melts into a different picture.

Image-to-video prompting is a different discipline from text-to-video. With text-to-video you are describing a world from scratch. With image-to-video, the world already exists — your job is to describe motion, not subject. Get that distinction wrong and you fight the model the entire time. Get it right and a single good still becomes a clean, controllable clip.

This guide covers how to write image-to-video prompts that respect your source frame across Kling, Runway, Luma, and Pika, with the phrasing each model rewards and copy-paste templates you can adapt.

Why Image-to-Video Needs Different Prompts

In text-to-video, your prompt has to invent the subject, the environment, the framing, and the motion. In image-to-video, three of those four are already decided by the image. If you re-describe the subject in detail, you are giving the model permission to reinterpret it — which is exactly how faces drift, products warp, and backgrounds reshuffle between frames.

The mental model that works: the image is the noun, the prompt is the verb. Your prompt should be almost entirely about what moves, how fast, and in what direction. Anything you say about the subject’s appearance should only reinforce what is already in the frame, never add new detail.

A practical rule from testing across models: if a clip features a specific person, product, or brand, you should almost always start from an image rather than a pure text prompt. Text-to-video is for invented worlds; image-to-video is for protecting something that already looks right.

The Four Things Your Prompt Should Describe

For any image-to-video prompt, focus the word budget on these:

  1. Primary motion — what the main subject does (turns head, takes a step, blinks, smiles, raises a cup).
  2. Secondary motion — ambient life that sells realism (hair moving, steam rising, fabric settling, leaves shifting, light flickering).
  3. Camera behavior — does the camera hold, push in slowly, drift, or orbit? Image-to-video often looks best with the camera nearly still.
  4. Pace — slow and subtle almost always beats fast and dramatic, because aggressive motion is where coherence breaks.

Notice what is missing: subject description. You do not re-describe the person’s clothing or the product’s shape. The image already carries that.

Subject holds the pose from the image. She slowly turns her head to look off-frame left, a small smile forming. Her hair shifts gently. Camera holds completely still. Subtle, slow motion.

That prompt animates a portrait without inviting the model to redesign the face.

Slow Beats Fast, Every Time

The single most common image-to-video mistake is asking for too much motion. A still has one frame of truth; every frame after it is the model’s guess, and the further the guess travels from the source, the more likely it warps.

Subtle motion keeps the clip anchored to the image. Instead of “she runs across the room,” try “she takes one slow step forward.” Instead of “the car speeds down the highway,” try “the car rolls forward slowly as the camera holds.” You can always generate a second clip for the big move; you cannot un-melt a face.

Speed words that keep image-to-video stable: slowly, gently, subtly, gradually, drifts, settles, barely. Speed words that tend to break it: quickly, rapidly, suddenly, whips, races.

Model-by-Model Notes

The four image-to-video models reward slightly different phrasing.

Kling

Kling handles complex human motion from a reference still better than most — body mechanics, weight shifts, and gestures come through cleanly. It responds well to plain, physical descriptions of movement and has solid camera-movement control when you want directed motion.

The man in the image takes a slow breath, his shoulders rising and settling. He shifts his weight to his other foot. Camera pushes in very slightly. Natural, grounded motion.

Keep Kling prompts concrete and physical. Abstract mood words do less here than literal body actions.

Runway

Runway is the strongest all-rounder for directed motion and clean production controls — reference-image support and camera control make it the pick when you need a specific, repeatable camera move on top of your still.

Hold the composition from the image. Slow dolly-in toward the subject. Subject’s expression softens, eyes blink once. Background remains static. Smooth, controlled motion.

Runway rewards naming the camera move explicitly. If you want a push-in, say push-in; do not leave the camera to chance.

Luma

Luma’s image-to-video leans dreamlike and fluid — it is the right choice for abstract, narrative, or music-video sequences where smooth, slightly surreal motion is a feature, not a bug.

Gentle drifting motion through the scene in the image. Atmospheric particles float in the air. Soft, flowing camera movement. Ethereal, continuous motion.

Lean into Luma’s strengths. Ask for flow and atmosphere rather than precise mechanical action.

Pika

Pika is best for fast, simple creative experiments, and its first-and-last-frame feature makes it strong for transition-style clips where you define a start image and an end image and let it interpolate between them.

Animate from the source image with gentle motion. Subtle camera drift forward. Light shifts softly across the subject. Clean, simple movement.

Use Pika when you want quick iterations or a frame-to-frame morph rather than a long controlled take.

When the Subject Drifts: Fixing Identity Loss

The most painful image-to-video failure is identity drift — the face or product slowly becoming a different one. Four fixes that reliably help:

  • Reduce motion. Less movement means fewer frames where the model can wander. Dial the action back before anything else.
  • Anchor with one identity cue. A single short phrase that matches the image (“same face, same outfit”) can help, but more than one invites reinterpretation.
  • Keep the camera still. Camera motion plus subject motion compounds drift. Hold the camera and let only the subject move.
  • Shorten the clip. Generate 4-5 seconds instead of 10. Drift accelerates over time; shorter clips stay truer to the source.

If a specific brand asset or face has to stay perfect, the consistency techniques in our consistent characters guide carry over directly to image-to-video work.

If you want clean image-to-video prompts without hand-tuning phrasing for each model, LzyPrompt turns one motion brief into model-tuned prompts for Kling, Runway, Luma, and Pika — so you spend your time choosing the best take, not rewriting the same idea four times.

Copy-Paste Templates

Adapt these by swapping the [bracketed] parts. Each assumes you have already uploaded a strong still.

Portrait, subtle life:

Subject holds the pose from the image. [Small action: slight smile / single blink / slow head turn]. Hair and clothing move gently. Camera holds still. Slow, subtle motion only.

Product, hero rotation:

The product from the image rotates slowly in place. Soft studio light glints across the surface as it turns. Camera holds still, shallow depth of field. Smooth, controlled motion.

Landscape / environment, atmospheric:

Bring the scene in the image to life with ambient motion. [Clouds drift / water ripples / grass sways / steam rises]. Camera drifts forward very slowly. Continuous, gentle movement.

Character, single deliberate action:

The character from the image performs one action: [takes a slow step / raises a hand / looks up]. Everything else stays anchored to the source frame. Camera nearly still. Grounded, slow motion.

Cinematic push-in:

Hold the composition from the image. Slow push-in toward the subject. Minimal subject motion — only [a breath / a blink]. Background static. Smooth, deliberate.

Transition (first-to-last frame, Pika/Luma):

Animate from the source image toward [the second frame / a slightly different state]. Smooth interpolation, gentle camera drift. Continuous motion with no sudden jumps.

These work as written. For more aggressive or complex motion, generate in shorter chunks and stitch in your editor rather than asking one prompt to do everything.

Common Image-to-Video Mistakes

Re-describing the subject. Every detail you add about appearance is an invitation to reinterpret the image. Describe motion, not the noun.

Asking for too much motion. Big actions break coherence. Start subtle, add a second clip for the big move.

Camera and subject both moving fast. The two compound and accelerate drift. Pin one of them down.

Ignoring secondary motion. A perfectly still subject with zero ambient motion looks frozen and uncanny. A little hair movement or rising steam sells the shot.

Generating 10-second clips by default. Drift grows with length. Shorter clips stay truer to your source — generate long only when the motion is genuinely minimal.

FAQ

Which model is best for image-to-video in 2026?

It depends on the shot. Runway is strongest for directed camera motion and clean production control. Kling is excellent for human motion and realism from a reference still. Luma suits dreamlike, flowing, abstract sequences. Pika is best for fast experiments and first-to-last-frame transitions. Run the same still through a couple of them and keep the usable take, not the flashiest demo.

How much should I describe the subject in an image-to-video prompt?

As little as possible. The image already defines the subject. Spend your prompt on motion, camera behavior, and pace. Any subject detail should only reinforce what is already in the frame.

Why does my subject’s face change over the clip?

That is identity drift, and it usually comes from too much motion or a clip that runs too long. Reduce the action, hold the camera still, and generate a shorter clip. The fewer frames the model has to invent, the truer it stays to your still.

Can I add camera movement to image-to-video?

Yes, but use it sparingly. A slow push-in or gentle drift looks great; aggressive moves compound with subject motion and cause warping. Runway and Kling handle directed camera moves most reliably. Our camera movement prompts guide covers the phrasing each model responds to.

Should I use text-to-video or image-to-video?

Use image-to-video whenever a specific person, product, or brand asset has to stay accurate — start from a still you control. Use text-to-video for invented scenes where the model has freedom to design everything from scratch.


Image-to-video is the most reliable way to get a specific look on screen, because you start from something that already looks right. The discipline is restraint: describe the motion, protect the frame, and let the still do the heavy lifting.

LzyPrompt generates motion-focused image-to-video prompts for Kling, Runway, Luma, and Pika from a one-line brief. Generate your first prompt free, no credit card.

Bank K.

Bank K.

Founder, LzyPrompt

Builder of LzyPrompt. Creates AI video prompts to help content creators save time generating professional videos for YouTube Shorts and Facebook Reels.

@ifourth on X

Ready to Try LzyPrompt?

Create professional AI video prompts in seconds. Start your free trial today.

Start Free Trial

© 2026 LzyPrompt.com by 3AM SaaS OÜ | All rights reserved | Secure login via Beag.io