Scene segmentation field guide

Turn one video clip into scene-by-scene AI prompts.

Find meaningful scene boundaries, write one prompt card per visual job, and carry subject, motion, and light from one generated scene into the next.

Scene prompt workflowPublished 10 min read

The short answer

Split by visual decisions, not by every cut

A scene-by-scene video prompt is a sequence of connected production briefs. Each card has one visual purpose, one dominant action, and a defined handoff to the next card. The goal is not to transcribe every frame. It is to preserve the decisions that make the sequence readable.

Begin with a timecoded shot list, then group adjacent shots that serve the same action. A hard cut may stay inside one prompt; a continuous shot may need two prompts when its visual objective changes.

01

Shot

One camera view between edits. Useful evidence, but often too granular for one prompt.

02

Scene

A continuous visual event with one location, state, and dramatic job.

03

Prompt card

The generation brief for that scene, including its incoming and outgoing state.

Boundary test

Decide where a new prompt begins

Location or time changes

Split

The environment and light need a new visual setup.

The subject enters a new state

Usually split

The next prompt needs a clear starting condition.

Camera job changes

Split

A reveal, pursuit, and proof shot have different objectives.

A cut continues the same action

Usually combine

Multiple angles can still belong to one generation beat.

Only the crop gets tighter

Combine or split

Split only if the close-up reveals new information.

Six-pass workflow

Build the sequence before polishing the prose

  1. 01

    Mark observable changes

    Watch once for the story, then again for changes in location, time, subject state, camera job, and visual emphasis. Record timestamps without writing prompts yet.

  2. 02

    Group cuts into scene jobs

    Label each group as hook, setup, action, reveal, proof, transition, or payoff. Merge neighboring shots when they advance the same job.

  3. 03

    Extract only reusable decisions

    Record framing, camera motion, light, palette, texture, subject motion, and transition direction. Replace the reference premise instead of copying its characters or outcome.

  4. 04

    Write the start and end first

    Define what is visible at the first frame and what must be true at the last frame. The action between them becomes much easier to specify.

  5. 05

    Add a continuity ledger

    Keep a short shared list for identity, materials, wardrobe, palette, light direction, screen direction, and scale. Change a ledger item only on purpose.

  6. 06

    Review adjacent cards in pairs

    Compare scene 1 with 2, then 2 with 3. Check that every end state can plausibly become the next start state before generating the full sequence.

If the source is a Short, use the tighter hook–build–turn–payoff structure in the YouTube Shorts prompt workflow. The boundary method stays the same; only the pacing gets denser.

Prompt card anatomy

Give every scene the same six fields

Purpose

What this scene must communicate or change.

Start state

The visible condition inherited from the previous scene.

Subject action

One dominant action, written in observable terms.

Camera

Shot size, angle, movement, direction, and framing.

Look

Environment, light direction, palette, texture, and atmosphere.

End state

The final composition or motion passed into the next scene.

[Scene purpose and time range]
[start state] → [one subject action]
[shot size, angle, camera motion, screen direction]
[environment, lighting, palette, texture]
[end state and continuity handoff]

Worked sequence

One 12-second clip, four connected prompt cards

This original paper-boat example keeps one red focal point and a consistent screen direction. Each scene changes the visual job while its handoff supplies the next scene’s starting state.

01Question
00:00–00:02

Extreme macro of a red paper boat pinned against a rain-dark curb, water pressing against its folded edge, locked camera at gutter level, cool overcast light with one red focal point, the boat pointing screen right, end with the current beginning to lift its bow.

Handoff

Red boat · screen-right direction · bow begins to lift

02Release
00:02–00:05

Low close tracking shot as the same red paper boat breaks free and accelerates screen right through the gutter stream, wet asphalt and silver reflections sliding behind it, camera matching the boat’s speed, maintain cool light and scale, end as it approaches a circular drain opening.

Handoff

Same boat and scale · rightward speed · drain enters frame

03Turn
00:05–00:09

Camera arcs above the red paper boat as it circles the drain instead of falling in, water forming a clean spiral around it, cool street reflections shifting toward warm amber from below, preserve clockwise motion and the boat’s folded silhouette, end directly overhead.

Handoff

Clockwise spiral · overhead composition · amber appears

04Payoff
00:09–00:12

Seamless overhead pullback reveals the gutter spiral as part of a luminous street-map pattern, the red paper boat at its center, rain easing to small ripples, balanced cool asphalt and amber lines, slow final settle with open space above the boat for a closing caption.

Handoff

Resolve motion · preserve red focal point · leave caption space

Sequence review

Catch five common failures before generation

  • Turning every edit into a separate prompt, even when the visual job is unchanged.
  • Describing the middle of a scene but omitting its start and end states.
  • Repeating a global style phrase while camera direction or subject scale resets.
  • Packing dialogue, edit notes, music cues, and image instructions into one dense block.
  • Copying the reference subject and payoff instead of rebuilding the sequence around a new premise.

Once the cards connect cleanly, use the broader AI video recreation workflow to plan generation passes, compare outputs, and assemble the edit.

Practical questions

Scene-by-scene prompt FAQ

What makes two moments separate scenes in a video prompt?

Start a new scene when the location, time, subject state, camera job, or visual objective changes. A simple cut does not always require a new prompt if both shots continue the same action and share the same generation goal.

How many scene prompts should I create from one clip?

Use the fewest prompts that preserve meaningful visual changes. Combine micro-cuts that serve one action, and split only when a new scene needs its own subject state, camera instruction, or continuity handoff.

How do I keep AI-generated scenes visually consistent?

Carry a small continuity ledger between prompts: subject identity, wardrobe or material, palette, light direction, screen direction, scale, and the exact end state that becomes the next scene’s start state.