The 3-Step AI
Video Workflow
The exact Image → Upscale → Video pipeline I use to turn one photo of your face into a believable AI clone that actually moves — with the prompt templates and the rules that keep it from melting.
Want to skip the learning curve
and learn directly from me?
See if you qualify to work 1:1 with me — I'll build your AI clone and your content system with you.
Apply to work with me →Image first, then upscale, then video. Every stage feeds the next — skip one and the errors compound. Fix the face before you ever animate it.
Every prompt skeleton below is ready to paste. Keep the structure and change only the identity and the scene so it matches YOU.
One camera move. One main action. Explicit constraints on identity and warping. That single discipline is what separates a stable clip from a melting one.
Create the Base Image
Flux 2.0 / Seedream 4.5 / NanoBanana PRO
These rules are the difference between an "AI face" and a real person:
Set it up like a real photo (not "art"): lens, depth of field, studio lighting, RAW look. (Example: 85mm f/1.4, shallow DOF, studio lighting.)DFY prompt
Force real skin texture: visible pores, micro-wrinkles, slight asymmetry, freckles/skin marks if consistent.DFY prompt
Block the "beauty filter" in the prompt: specify no makeup, no retouching, no smoothing.DFY prompt
Use soft, even lighting: if you start too dramatic, the AI often "paints" the skin instead of describing it.DFY prompt
No tag lists: write one natural paragraph in a photographic description style (fewer bugs, more coherence).DFY prompt
Use this as the skeleton and only swap identity/scene:
If your base image is already "too perfect," video will destroy it. Fix skin/eyes/detail first, then animate. If you do it backwards, you'll waste hours.
Upscale
Leonardo / others → I use LUPA.ai
Because I need an upscale that adds detail without changing identity (same face, same texture, same features). The goal isn't "prettier" — it's more real and more stable for image-to-video.
Light/moderate upscale > aggressive upscale: push too hard and skin turns "waxy," and it shows in video.
Check the eyes: if reflections/iris details change, Kling/Veo will "wobble" later.
Check micro-details: pores, beard stubble, hairline, eyebrows (if the upscaler invents them, identity shifts).
Create the Video
Kling / Sora / Veo — and how to adjust your prompt for each
Pros. Excellent controllable image-to-video, clear camera movement, great for talking motion and realistic performance if you direct it well.
Cons. If you ask for too many things at once, it gets unstable. With Kling 2.6, you must be highly structured and usually positive-only.
Pros. Great when you want an engineered description (separate fields: camera, lighting, motion, ending, etc.).
Cons. If you write it too generic, you'll get pretty videos but not always controlled.
Pros. Often very strong cinematic quality and scene coherence, great for story-like sequences.
Cons. Less predictable if you want ultra-specific motion like "do exactly this micro-action" (depends on the shot).
Don't describe the photo. Describe only: camera + behavior + micro-actions + atmosphere + constraints. Mandatory structure: Scene → Camera → Subject → Action → Audio → Style → Constraints.
Fill each field, keep it structured, stay positive-only:
Use JSON with clear fields. The advantage is you separate what happens from how it's filmed.
One field per concern — describe what happens separately from how it's filmed:
Write like a director: one shot, clear context, readable action, simple camera, and few simultaneous requests. If you try to write a technical manual like Kling 2.6, it's often not ideal — go for cinematic narration instead.
The one rule for stable video
Kling states this openly — otherwise it collapses.