Image-to-Video Prompts: A Practical Motion Guide

Write image-to-video prompts that separate subject and camera motion. Use original examples, source-frame checks and a focused workflow for revising each shot.

By Muse Video

An image-to-video prompt should explain what changes after the starting frame. The image already supplies visual information. Your writing can concentrate on subject movement, environmental movement and camera behavior, then name the important details you want to remain consistent.

Google's video generation best practices recommend this motion-focused approach for its image-to-video workflow. Use it as a baseline to test in your chosen model, not as a guarantee of perfect preservation. The examples below are original shot briefs; they are not documented model outputs.

If you have not decided whether to use an image at all, first read text-to-video versus image-to-video. This guide begins after you have a source frame and want to write the motion.

Write the missing information

A photograph of a person by a window already shows a face, pose, light source and composition. Repeating every visible detail can make the prompt longer without deciding what should happen. The useful missing information might be a blink, a small head movement, a moving curtain or a slow camera approach.

Start by writing one sentence for the main change. Then decide whether the camera moves. Finally, add a short stability instruction for the parts that matter most to the project.

A reusable planning structure is:

Starting from the supplied image, [describe one visible action]. The camera [stays locked or follows one clear movement]. Keep [specific important details] consistent. The shot [continues without a cut or follows another explicitly planned transition].

Treat those brackets as decisions, not a magic formula. The result still depends on the input, model and supported controls. If the first test is confusing, the structure helps you identify which sentence to revise rather than appending more unrelated instructions.

Check the frame before you animate it

The source frame determines what is already visible and what a requested movement would reveal. Review it before writing a complex action.

Framing: Is the subject already cropped tightly? A push-in may hide a face or product detail. A wide gesture may extend beyond the original frame. Choose a crop that leaves room for the intended action.

Clarity: Can you see the detail you expect the model to preserve? A blurred logo, hidden hand or compressed face gives an uncertain starting point. Google's best-practice page specifically emphasizes a clear, well-composed source image.

Depth and occlusion: A flat photograph does not show every side of an object. A large orbit may require new geometry. If the unseen surface matters, investigate supported reference controls or plan a smaller camera move first.

Provider requirements: Check file formats, image sizing, aspect ratios and input roles in the selected tool. Google's image input documentation describes its accepted formats and possible resizing or cropping. Do not assume that an image accepted by one interface is accepted by every API.

Intent: Decide whether you are animating the existing composition or asking for a redesign. Changing the outfit, background, light and viewpoint together is a different job from adding one small movement to the original frame.

Keep three kinds of motion separate

Subject motion

This is the action performed by the main subject: blinking, turning slightly, walking, or moving an object. Name the visible event and its scale. “A small nod once, then a return to the starting pose” gives you a clearer test than “make the person dynamic.”

When continuity matters, start with movement that fits the existing pose. A subtle breath in a seated portrait introduces fewer new decisions than a jump into a different body position. That is a planning choice, not a claim that larger changes are impossible.

Environmental motion

This includes steam, leaves, reflections, mist, curtains or other surrounding elements. It can give a still image a sense of movement while the main subject remains stationary.

Pick one environmental action first. If leaves, clouds, water, light and clothing all move strongly at once, it becomes harder to tell which instruction changed the result. You can add layers after deciding which movement the scene actually needs.

Camera motion

A push-in, pan, tilt or orbit changes how the viewer sees the scene. Describe it independently of subject action. A locked camera with moving steam is a different request from a camera approaching a still cup.

Google's prompt guide and Adobe's video prompting guidance describe camera vocabulary. Availability and reliability remain model-specific. Use one clear movement and review how the model interprets it.

Five original motion prompts to adapt

Each example assumes that you supply the appropriate source image to the generation tool. The writing describes intent, not a promise that a face, label or object will remain exact.

1. A small portrait movement

Use the supplied portrait as the starting composition. The person makes one gentle blink and a small natural breath while looking just past the camera. Keep the camera locked, the face position stable and the window light unchanged. A few loose strands of hair move slightly.

This gives you one modest human action and a stable view. Before adding speech or a head turn, check the actual output for facial continuity and unwanted changes. You can customize the portrait example in the library.

2. A stationary product with a shallow camera move

Starting from the supplied bottle photograph, move the camera through a short, shallow arc around the front. The bottle stays upright and stationary. A soft highlight moves across the glass. Keep the cap, label placement and backdrop consistent, and avoid a cut to a new view.

A shallow arc is a restrained first question to test. A much wider orbit asks the model to reveal surfaces not shown in the original photograph. If precise packaging text is essential, review it separately and plan an editing step rather than assuming the prompt guarantees it.

3. Environmental movement in a landscape

Use the supplied forest frame. The camera remains locked while a few leaves sway gently and a thin patch of mist drifts between the trees. Keep the main trunks, horizon and sunlight direction stable. The movement stays quiet and continuous.

If the result changes too much, test leaves and mist in separate attempts. This distinguishes whether one environmental instruction is doing useful work or making the image harder to control. The forest example provides another starting brief.

4. A small action in a food image

Begin from the supplied salad photograph. A few herb leaves fall gently onto the center of the plate while the dish and camera stay still. Keep the ingredient arrangement, table color and lighting consistent. Show one finishing action with no change of scene.

The prompt adds one event to the existing composition. If it instead asks for the ingredients to be chopped, mixed, plated and served, it becomes a sequence rather than a small animation. Plan those beats as separate shots when that makes the project easier to control.

5. A quiet interior

Use the supplied room image as the starting frame. The camera remains locked. A light curtain moves gently near the open window while the furniture, wall lines and floor stay stable. The daylight stays soft and consistent throughout the shot.

Straight architectural lines give you useful details to inspect. If furniture changes or the room appears to bend, simplify the movement and check the input quality before adding a camera move. The intended action is the curtain, not a redesign of the room.

Use preservation and exclusion notes carefully

Preservation notes should be specific enough to review. “Keep everything perfect” is difficult to act on. “Keep the bottle upright and the label in the same position” identifies visible criteria, while still requiring inspection of the actual clip.

Do not mix preservation with contradictory redesign instructions. Asking to keep the source lighting while adding a new strong light from the opposite side leaves an unresolved choice. Decide which outcome matters, or test the lighting change in a separate attempt.

Keep negative-prompt fields separate from the main motion brief when the provider supports them. Google's prompt guide recommends describing unwanted elements in its negative field rather than using instructive phrases such as “don't.” Another model may expose a different mechanism. Copy your exclusions into the supported control instead of assuming a universal syntax.

Duration, aspect ratio and audio also need a separate check. Select supported values in the interface or API. If sound is part of the project, confirm that the chosen model supports the intended audio workflow; an image and a sentence about sound do not establish that capability.

Troubleshoot with a controlled sequence

When a clip misses the brief, first describe the visible mismatch in plain language. “The camera moved left even though I wanted it locked” suggests a more focused revision than “the result was bad.”

MismatchA useful next test
The camera driftsRemove other camera terms, ask for a locked frame and keep subject motion modest.
The subject changes too muchReview source clarity and test a smaller action before changing the scene.
A product surface is inventedReduce the viewpoint change or investigate supported additional references.
Too many elements animateKeep only the main motion instruction, then add other layers one by one.
Motion starts outside the cropReframe the input or choose an action that fits the visible area.
A label or face is inaccurateReview it as a separate acceptance criterion; change the workflow if exact reproduction is essential.

Keep the source file, provider, model and output settings fixed while trying a specific change. Save the complete prompt and a short observation beside each result. The general prompt guide includes a comparison-log format you can reuse.

This process can identify useful revisions, but it is not evidence that wording alone can solve every failure. Some requirements may need editing, a different input, supported reference controls or another model. The point of a clear brief is to make that decision easier.

Use the builder as a planning tool

  1. Select Image or Reference in the Muse Video prompt builder, then describe the motion you want. Reference image bytes stay in the browser for planning.
  2. Set the camera, motion level and intended format. Review the composed brief and copy it as plain text.
  3. Upload the actual source file separately to your generation provider and select its supported settings. A local planning reference is not automatically sent to that service.
  4. Review the returned clip against your chosen criteria. The sample footage shown by Muse Video is a separate player demonstration and does not prove that the prompt produced an output.

If you need more starting scenes, use the example library. If you are choosing an input mode rather than writing its motion, return to the workflow comparison.

Sources and scope

Sources checked on September 25, 2026: Google's generation best practices, video prompt guide and image input documentation, plus Adobe's video prompting guidance. The five example briefs and troubleshooting sequence are original editorial suggestions. No model benchmark or first-hand generation result is claimed.

Questions about this workflow

Should I describe everything in the source image?
Start with the movement the image does not show. Google's image-to-video guidance recommends focusing on motion rather than repeating the subject, scene and style already in the frame. Adapt that starting point to your provider.
How do I keep the camera still?
State that the camera is locked and the framing stays unchanged, then describe the subject or environmental movement separately. Remove conflicting requests for a pan, zoom, orbit or viewpoint change and inspect the returned clip.
Can I use several reference images?
Only if the selected model supports the relevant reference input. First frames, last frames and reference images can have different roles. Label the purpose of each asset and check the provider's actual controls.

Put the guide into practice

Choose an example or build your own shot brief. Copy the prompt into the video model you use.

Explore AI models on OmniAKey
Image-to-Video Prompts: A Practical Motion Guide · Muse Video