TubePrompter Learn

Turn a reference video into an image-to-video prompt

Turn a reference shot into two separate inputs: an image showing your new scene at the start, and a prompt describing what happens next. This guide takes you from video analysis notes to an image-to-video attempt you can review and revise.

Image-to-video workflowPublished 8 min read

1. Give the image and the text different jobs

A reference analysis may describe the wardrobe, light, camera angle, action, and ending in one detailed record. Keep that record as your brief. For the image-to-video step, decide which information should already be visible and which information describes a change.

Starting image: what exists now

Your subject, starting pose, surroundings, framing, and visible appearance. Prepare a real image file that shows those decisions.

Motion prompt: what changes next

The action, movement direction, camera behavior, and ending you want to inspect. Describe the change from the image you actually supplied.

Runway's current prompting guide explains that an image establishes the visual setup while the text guides movement. Its advice also allows visual detail when introducing something new or describing a transformation. You do not need to strip out every noun; you need to avoid restating the whole image.

If you are still mapping the reference, begin with the scene-by-scene prompt workflow. Here, work on one selected shot rather than pasting a full sequence into one generation.

2. Prepare a starting image that can support the action

  1. Write the opening state. Note who or what is visible and where the action begins. For a drawer opening, start with the drawer closed; for a pickup, start with the object within reach.
  2. Choose your own visual content. Carry over a useful composition or movement idea, then design your own subject and setting. Prepare an original photograph, illustration, or generated still for the new shot.
  3. Inspect the actual image. Check object edges, hands, faces, and the space needed for the action. A prompt describing a perfect image does not repair a flawed file automatically.
  4. Check the selected generation mode. Use an image-to-video workflow that accepts a starting image. Follow that tool's current input and duration requirements; a text description is not an image attachment.

Google's video generation best practices similarly recommend a clear source image and motion-focused text. They distinguish camera movement, subject animation, and environmental changes. Choose which of those carries your shot before adding secondary movement.

Save palette and material decisions in your image brief using the color and texture card. Keep a copy of the approved starting image so later attempts can use the same input.

3. Write one movement with a visible finish

Draft the action in plain language, then read it while looking at your image. Can you point to the subject that moves? Is there room for the camera view you want? What would count as a completed action? Use the camera movement guide if you need to separate a moving subject from a moving viewpoint.

Starting image: [approved image file; keep this in your notes]
Subject action: [who moves + direction + pace]
Camera: [one movement or a fixed position]
Environment: [optional supporting motion]
Ending: [the visible state to reach]
Review: [one action check + one continuity check]

This is a planning card, not syntax to paste into every model. Combine the relevant motion fields into a short prompt. Keep the filename and review notes outside the prompt. A camera can move while a subject stays still, so state those choices separately.

Prefer a concrete endpoint such as a packet held at chest height to a mood instruction such as an inspiring reveal. If the shot needs several distinct events, first try the one event that makes its purpose clear. Add the others only after reviewing that attempt.

4. Try three different image and motion pairings

These original examples are illustrative briefs, not tested generation results. Prepare the described image first, then adapt the motion text to the tool you use.

Product reveal: move the camera

Prepare this image: An original still of a folded olive-green travel bag on a pale bench. The bag sits left of center, with clear space to its right. Soft side light reveals the seams.

Motion prompt

The camera slides slowly to the right past the bag, revealing its side pocket. The bag remains resting on the bench. The movement settles into a close view of the pocket.

Inspect and revise: At the end, can you see the side pocket rather than a newly invented front face? Does the bag stay in contact with the bench? If the reveal fails, begin with an image that already shows part of the pocket.

Character beat: move one subject

Prepare this image: An original still of a gardener beside a small potting table. Both hands are visible, with one hand resting beside a seed packet. The camera view includes the packet and the gardener’s face.

Motion prompt

With the camera fixed, the gardener lifts the seed packet to chest height, glances at it, then smiles. The packet stays in the same hand throughout the action.

Inspect and revise: Review the pickup, the middle of the lift, and the final hold. If the packet changes hands or shape, simplify the next attempt to the lift alone before adding the expression.

Atmosphere: move the environment

Prepare this image: An original still of a bicycle leaning against a stone wall, with a shallow puddle in the foreground and fallen leaves along its edge.

Motion prompt

The camera holds its position. A light breeze carries two leaves across the puddle, making small ripples. The bicycle remains leaning against the wall as the ripples spread and soften.

Inspect and revise: Look at the bicycle as well as the leaves. If the bicycle starts rolling, reduce the environmental action to the ripples first and inspect whether the starting image suggests instability.

5. Review the handoff from still image to moving shot

Watch once at normal speed, then inspect the beginning, middle, and end. Record whether the action completed and whether the important object stayed recognizable. A pleasing first frame alone does not answer either question. Compare against your new brief rather than expecting an exact reconstruction of the reference video.

The clip barely moves
Identify one visible action and remove repeated appearance notes. Check that the action can happen from the starting pose.
The wrong object moves
Identify it by position or a short visible feature, such as the person beside the table. Give the other subject a clear resting state.
The scene changes unexpectedly
Check for a hidden request for a second location or shot. Keep one continuous action, then plan the next shot separately.
The output begins with the action already finished
Revise the starting image so it shows the setup: hand beside the packet, closed drawer, or unturned object.
The ending never arrives
Shorten the action or use a longer duration if the selected tool supports it. Remove intermediate events before adding more timing instructions.

Save the image, prompt, model or mode, and generation settings with each attempt. Change either the image or the motion wording first so you can track your revision. Reusing the same inputs does not guarantee identical outputs; the log helps you compare decisions, not prove that one wording always performs better.

Use TubePrompter to build the reference brief

Sign in and analyze a supported YouTube video in TubePrompter. Review the scene descriptions and start-frame and end-frame prompt text against the reference. Analysis detail depends on your plan. Those frame prompts are descriptions to adapt, not recovered original prompts or actual image files.

Use the start description to plan your own still, and use the end description to decide where its motion should lead. The preparation and comparison steps above are a manual workflow alongside the analysis. The YouTube extraction walkthrough explains how to begin with a reference link.

Analyze a YouTube reference

Image-to-video prompt FAQ

How do I turn a reference video into an image-to-video prompt?

Choose one shot, separate its opening appearance from the movement that follows, and prepare an original starting image. Write a short motion prompt describing the subject action, camera behavior, and intended ending. Compare the generated beginning, middle, and end with your written plan.

Are a start-frame prompt and a starting image the same thing?

No. A start-frame prompt is text describing an image. An image-to-video tool needs an actual image in its image input. Use the description to prepare that image, check it, and then write the motion instructions separately.

Should I paste the entire video analysis into the motion prompt?

Use the analysis as editing notes. In the Runway and Google image-to-video workflows cited here, the image supplies the visual starting point and the text primarily directs motion. Keep only the details needed to identify the moving subject or explain a deliberate change.