-
Step 1: Build the character list first
What to do: Begin by defining the characters before generating scenes. Keep each character’s identity and role available as the reference point for the rest of the workflow.
What success looks like: Every planned scene has a clear character identity to use.
Common mistake to avoid: Do not begin with disconnected shots that have no shared character reference.
-
Step 2: Generate the character reference
What to do: Use Nano Banana 2 to create the visual foundation for the character. For stronger Nano Banana 2 image generation, describe the identity, appearance, expression, and intended visual style in a focused prompt.
What success looks like: The reference clearly communicates who the character is and how the character should appear.
Common mistake to avoid: Avoid changing several defining traits between reference generations.
-
Step 3: Plan the scene around the same identity
What to do: Move from the character list to scenes, then define the action, emotion, setting, lighting, and camera direction. This is where cinematic AI video benefits from a stable character foundation.
What success looks like: The scene description adds new action and atmosphere without replacing the character’s core identity.
Common mistake to avoid: Do not use vague prompts that leave expression, framing, or performance undefined.
-
Step 4: Add multimodal references when more control is needed
What to do: Use Smart Reference to provide images, videos, and audio as references. Multiple input types can help communicate visual direction, motion expectations, or performance context.
What success looks like: The generated scene reflects the intended reference material while preserving the character’s recognizable features.
Common mistake to avoid: Avoid adding references that conflict with one another in style or subject.
-
Step 5: Choose Director Mode or Instant Mode
What to do: Use Director Mode to start with a storyboard and render with more control over each shot. Use Instant Mode to go directly from prompt to video when speed is the priority; its duration is up to 30 seconds.
What success looks like: Your mode matches the project: controlled multi-shot planning or fast prompt-to-video creation.
Common mistake to avoid: Do not expect Instant Mode to provide the same shot-by-shot planning as Director Mode.
-
Step 6: Generate clips and composite the final video
What to do: Render the video clips, inspect continuity, and assemble the final sequence. Depending on the workflow, duration can range from 30 seconds to 10 minutes, while available video models include Kling 3.0, Seedance 2.0, MiniMax H3, and Seedance 2.5.
What success looks like: The final sequence feels like one connected story rather than unrelated generations.
Common mistake to avoid: Do not export before checking transitions, character identity, voice, music, effects, and pacing.