nicola.ai
● Free resource — 3-step workflow

The 3-Step AI
Video Workflow

The exact Image → Upscale → Video pipeline I use to turn one photo of your face into a believable AI clone that actually moves — with the prompt templates and the rules that keep it from melting.

Work with me

Want to skip the learning curve
and learn directly from me?

See if you qualify to work 1:1 with me — I'll build your AI clone and your content system with you.

Apply to work with me →
Instagram — @nicola.ai
How to use it
STEP 1
Work the pipeline in order

Image first, then upscale, then video. Every stage feeds the next — skip one and the errors compound. Fix the face before you ever animate it.

STEP 2
Copy the templates, swap the identity

Every prompt skeleton below is ready to paste. Keep the structure and change only the identity and the scene so it matches YOU.

STEP 3
Obey the golden rules

One camera move. One main action. Explicit constraints on identity and warping. That single discipline is what separates a stable clip from a melting one.

The workflow — Image → Upscale → Video
01

Create the Base Image

Flux 2.0 / Seedream 4.5 / NanoBanana PRO

Tools mentioned
Flux 2.0Great for clean photorealism and strong composition, but it often "beautifies" the face unless you hard-lock it with strict rules.
Seedream 4.5Strong aesthetics and consistency, but skin can look "too perfect" unless you force real texture/imperfections.
NanoBanana PRO (my choice)I use it when the priority is a believable face + real texture + photographic control, and I want an image that already looks like a real photoshoot without needing heavy post. The advantage is you can write editorial-photography style prompts and get a more stable base for animation (less "melting face" once you go to video).
Non-negotiable rules for a super realistic face (with visible imperfections)

These rules are the difference between an "AI face" and a real person:

01

Set it up like a real photo (not "art"): lens, depth of field, studio lighting, RAW look. (Example: 85mm f/1.4, shallow DOF, studio lighting.)DFY prompt

02

Force real skin texture: visible pores, micro-wrinkles, slight asymmetry, freckles/skin marks if consistent.DFY prompt

03

Block the "beauty filter" in the prompt: specify no makeup, no retouching, no smoothing.DFY prompt

04

Use soft, even lighting: if you start too dramatic, the AI often "paints" the skin instead of describing it.DFY prompt

05

No tag lists: write one natural paragraph in a photographic description style (fewer bugs, more coherence).DFY prompt

Practical template — NanoBanana PRO (photoreal base)

Use this as the skeleton and only swap identity/scene:

Front-facing close-up portrait, shot on an 85mm f/1.4 lens, soft even studio lighting, shallow depth of field, RAW photo look. Natural skin texture with visible pores, subtle freckles, faint expression lines, slight facial asymmetry, realistic hair flyaways. Neutral expression, no makeup, no retouching, no beautifying filter. High-detail photorealism.

Ruthless note

If your base image is already "too perfect," video will destroy it. Fix skin/eyes/detail first, then animate. If you do it backwards, you'll waste hours.

02

Upscale

Leonardo / others → I use LUPA.ai

Tools mentioned
Leonardo UpscalerFast and convenient, but it often "enhances" micro-details in a way that can make skin look plasticky if you're not careful.
Other upscalersUseful for sharpness, but some change the face too much (which is a disaster for an avatar).
Why I use LUPA.ai

Because I need an upscale that adds detail without changing identity (same face, same texture, same features). The goal isn't "prettier" — it's more real and more stable for image-to-video.

Upscale rules so you don't ruin the face
01

Light/moderate upscale > aggressive upscale: push too hard and skin turns "waxy," and it shows in video.

02

Check the eyes: if reflections/iris details change, Kling/Veo will "wobble" later.

03

Check micro-details: pores, beard stubble, hairline, eyebrows (if the upscaler invents them, identity shifts).

03

Create the Video

Kling / Sora / Veo — and how to adjust your prompt for each

Key differences
Klingespecially for avatars/creators

Pros. Excellent controllable image-to-video, clear camera movement, great for talking motion and realistic performance if you direct it well.

Cons. If you ask for too many things at once, it gets unstable. With Kling 2.6, you must be highly structured and usually positive-only.

Veodirector-level control via JSON

Pros. Great when you want an engineered description (separate fields: camera, lighting, motion, ending, etc.).

Cons. If you write it too generic, you'll get pretty videos but not always controlled.

Soramore "generative cinema," less rigid engineering

Pros. Often very strong cinematic quality and scene coherence, great for story-like sequences.

Cons. Less predictable if you want ultra-specific motion like "do exactly this micro-action" (depends on the shot).

Kling prompting approach (golden rule)

Don't describe the photo. Describe only: camera + behavior + micro-actions + atmosphere + constraints. Mandatory structure: Scene → Camera → Subject → Action → Audio → Style → Constraints.

Kling 2.6 template (stable and reliable)

Fill each field, keep it structured, stay positive-only:

Scene: environment, lighting, mood Camera: one camera move only (push-in OR arc OR handheld micro-sway) Subject: vibe and posture Action: 1 main action + 1–2 micro-events (no chaos) Audio: ON/OFF; if ON, exact line; minimal ambience Style: [your look — editorial / documentary / cinematic] Constraints: no warping, stable identity, stable colors, no text/logos

Veo prompting approach

Use JSON with clear fields. The advantage is you separate what happens from how it's filmed.

Veo — JSON fields

One field per concern — describe what happens separately from how it's filmed:

{ "description": "", "style": "", "camera": "", "lighting": "", "environment/room": "", "elements": "", "motion": "", "ending": "", "text": "" }

Sora prompting approach

Write like a director: one shot, clear context, readable action, simple camera, and few simultaneous requests. If you try to write a technical manual like Kling 2.6, it's often not ideal — go for cinematic narration instead.

Universal rule

The one rule for stable video

One camera move
One main action
Max 1–2 micro-events
Explicit constraints on identity / warping / text

Kling states this openly — otherwise it collapses.

nicola.aiThe 3-Step AI Video Workflow — free resource