ENTIRE GUIDE ON HOW TO CLONE YOURSELF WITH THE BEST QUALITY IN THE WORLD
Stop manufacturing every frame. Define the story, lock the visual DNA that matters, and let the video model construct the sequence — the full production manual, the end-to-end SOP, the templates and every checklist I use with my team.
Want to skip the learning curve and learn directly from me?
This is the exact system my team runs so a business owner never has to record. I'll build your AI clone and this production system with you, 1:1.
Apply to work with me → Instagram — @nicola.ai- 01Executive Overview and Mental Model
- 02Why the Old System Became a Bottleneck
- 03The New Storyline-First Architecture
- 04Storyline First: The Non-Negotiable Principle
- 05Turning an Idea into a Director-Level Storyline
- 06Subject Architecture and Character Logic
- 07Motion, Attitude and Performance Direction
- 08Location Logic and Spatial Consistency
- 09Object Logic and Prop Consistency
- 10Camera Direction, Framing and Transitions
- 11Reference Architecture
- 12Building and Approving Character Sheets
- 13Building and Approving Location References
- 14Building and Approving Object References
- 15Reference Numbering, File Naming and Upload Discipline
- 16Claude as the Translation Layer
- 17Splitting a Video into Seedance Generations
- 18Seedance 2.5 Execution Workflow
- 19Full End-to-End SOP
- 20Quality Control and Review Framework
- 21Troubleshooting and Failure Modes
- 22Complete Worked Example: Busy Business Owner
- 23Reusable Templates and Production Forms
- 24Final Checklists
- 25Operating Rules and Glossary
The most important change in this system is not a new tool. It is a different way of thinking about AI video production. The old process treated every finished image as a building block. The new process treats the storyline and the visual references as the building blocks.
That distinction changes the entire workflow. Instead of asking, “Which image do I need to generate next?”, you begin by asking, “What exactly happens in this scene, and which visual elements must remain consistent while it happens?”
The unit of work is no longer the individual image. The unit of work is a clearly directed scene supported by controlled references.
Read the mental model first
Chapters 1–4 explain why the system works. Skip them and every later step turns into guesswork.
Run the SOP, phase by phase
Chapter 19 is the operating procedure. Copy the templates in chapter 23 and fill them for your own video.
Review against the storyline
Judge every clip by whether it executed the scene, not by whether it looks cinematic. Chapters 20–21 tell you what to check and how to fix it.
The mental model
What changed, why the old way stopped scaling, and the one principle everything else depends on.
Executive Overview and Mental Model
Storyline, references, Claude and Seedance each do one job. Keep them separate and the system stays fast.
Storyline = the creative instruction layer.
The storyline describes everything that happens in the video in a way a director could understand. It is the bridge between your idea and the generation model.
Why it mattersIf the storyline is vague, every later step becomes guesswork. If the storyline is precise, reference creation and video prompting become dramatically easier.
ExampleInstead of writing “the owner wakes up stressed,” you describe the alarm, the missed reach, the spilled water, the second reach, the camera movement and the exhausted attitude.
Common mistakeTreating the storyline as a summary rather than a production document.
Approval testA stranger should be able to read the storyline and understand what the viewer will see, shot by shot.
References = the visual identity layer.
References tell the model what must stay visually controlled. They prevent the model from inventing a new face, a different room, or a different hero object every time the camera changes.
Why it mattersThe system only works efficiently if you control what matters and delegate what does not.
ExampleThe main owner gets a Character Sheet. A wife whose face is barely visible can simply be described.
Common mistakeCreating references for every incidental person, chair, cup and background object.
Approval testAsk: would I care if this element changed? If yes, reference it. If no, describe it and move on.
Claude = the translation layer.
Claude takes your human creative direction and converts it into a structured Seedance-ready prompt. It should not replace your creative decisions; it should encode them.
Why it mattersYou want the model generating the video to receive clear camera, action, reference and timing instructions without forcing you to hand-write every technical detail.
ExampleYou paste a storyline segment, upload the matching references, specify the intended duration and editing energy, and Claude produces the generation prompt.
Common mistakeGiving Claude an unclear concept and expecting it to solve story, continuity, timing and references all at once.
Approval testClaude's output must still match the storyline you approved.
Seedance 2.5 = the scene-construction layer.
Seedance is where the actual moving sequence is generated. It combines the prompt with the references and attempts to construct the requested action over time.
Why it mattersThis is what allows the workflow to stop depending on a manually generated starting image for every single shot.
ExampleA single generation can move from one camera framing to another, follow an action, or execute several beats that previously required separate image generations.
Common mistakeAssuming the model will understand unspoken creative decisions.
Approval testJudge the output by whether it executed the intended scene, not merely by whether it looks cinematic.
Why the Old System Became a Bottleneck
Complexity grew with every shot, angle, action and object — and the editor paid for it.
The old system was image-first. Once the idea was approved, production revolved around creating all the individual images required to manufacture the video. This worked, but complexity increased almost linearly with the number of shots, camera angles, actions and objects in the video.
Every new visual beat created a new task.
If the camera angle changed, if an assistant entered, if a dossier appeared, or if the subject changed position, you often needed another image. One creative idea could quickly become a long queue of image-generation tasks.
Why it mattersThe bottleneck was not imagination. The bottleneck was manufacturing every visual state manually.
ExampleA dynamic hook could require six or more distinct images before you even started animating.
Each image carried its own prompt and reference burden.
You had to tell the image model who the subject was, where they were, what they were wearing, what object was present, what the camera saw, and what style or lighting you wanted.
Why it mattersThe cost of one scene was multiplied by the number of still images required.
ExampleWhen a sequence required 20–30 images, that meant 20–30 opportunities for drift, wrong framing, wrong identity or inconsistent props.
Continuity had to be rebuilt repeatedly.
Even when the creative concept stayed the same, each still-image generation could reinterpret the face, body, outfit, room or object.
Why it mattersThe creator had to manually police continuity across dozens of generated assets.
ExampleA dossier could become a different size. A room could change architecture. A face could subtly drift between shots.
Editing became the glue holding the system together.
After creating the stills, you generated short video clips from them and then assembled those clips to simulate one continuous scene.
Why it mattersThe editor was compensating for fragmentation that originated earlier in the workflow.
ExampleThe more clips you needed, the more transitions and continuity problems had to be solved in post-production.
Trying to solve a structural production problem by simply generating more images.
If a scene needs many individual stills before it works, ask whether the new storyline-first method can consolidate them into one generation.
The old system was not wrong. It was simply optimized for a model era where the safest way to control a video was to control every starting frame.
The New Storyline-First Architecture
Idea, script, storyline, references, Claude, Seedance, assembly. Finish one before touching the next.
The new architecture separates the work into seven distinct layers. Keeping those layers separate reduces confusion and unnecessary rework.
Layer 1 — Idea
The business or creative concept. What is the point of the video? What should the viewer understand or feel?
Layer 2 — Script
The words: dialogue, voiceover, spoken lines or text that carries the message.
Layer 3 — Storyline
The visual story that happens around the script: actions, movements, emotions, location changes, objects and camera behavior.
Layer 4 — References
The assets that control the important visual DNA: subjects, locations and objects.
Layer 5 — Claude
The technical translation of the storyline into a structured prompt.
Layer 6 — Seedance 2.5
The generation of the moving sequence.
Layer 7 — Assembly
Combining several generations, audio, captions and editing into the final asset.
A clean separation prevents you from solving the wrong problem at the wrong stage.
Jumping ahead to generation while earlier creative decisions are still changing.
Before moving to the next layer, the current layer should be stable enough that later work will not be thrown away.
Storyline First: The Non-Negotiable Principle
Know what the scene is supposed to do before you generate anything.
The storyline is the single most important document in the new system. It can live in Notion, Google Docs or any other writing environment. The tool does not matter. The quality of the description does.
Before generating a Character Sheet, location, prop, or video, know what the scene is supposed to do.
The seven questions every storyline must answer
Who is present?
List the subjects that actually matter. Distinguish between controlled subjects and incidental people.
Why it mattersYou cannot plan references properly until you know who the recurring subjects are.
What does each subject do?
Describe physical actions in sequence. AI video responds better to observable behavior than to vague concepts.
Why it matters“He is stressed” is less useful than “he rubs his face, looks at the clock, exhales, and pushes the phone away.”
How do they feel?
Add attitude because the same action can look completely different depending on performance.
Why it mattersA calm reach toward an alarm is a different scene from an exhausted, irritated reach.
Where does it happen?
State the location and decide whether the space requires a controlled reference.
Why it mattersThe location determines background, lighting logic, props, spatial continuity and camera possibilities.
What objects matter?
Identify objects that carry the action or repeat across shots.
Why it mattersA hero dossier, alarm clock or specific car may require its own reference.
What does the camera do?
Describe shot type and movement only where it matters.
Why it mattersThe new system can handle switching angles, but it needs direction when the change is intentional.
How does the scene progress?
Write the beginning, middle and end of the generation in temporal order.
Why it mattersSeedance is generating time, not a static frame. Sequence is the core of the prompt.
Leaving this decision to the model when it directly affects the message.
The storyline should answer the question without forcing the operator to guess.
Writing the storyline
How to turn an idea into director-level direction: subjects, motion, attitude, locations, objects and camera.
Turning an Idea into a Director-Level Storyline
One sentence, then the script, then the subject map, then verbs the viewer can see.
5.1 Start with the one-sentence idea
The one-sentence idea keeps the storyline anchored. It should describe the situation and the transformation, not every shot.
A busy professional has no time or energy to record content. When his assistant asks him to film, he refuses. He then discovers an AI clone, allowing the team to create content without requiring him to record.
5.2 Write the script before over-designing visuals
The script tells you how much narrative space you have. If the spoken content lasts 25 seconds, your storyline must fit that reality. Visual ambition that ignores script timing creates clips that feel rushed or disconnected.
5.3 Build the subject map
Count the subjects.
Write down the number of meaningful subjects before writing detailed shots.
Why it mattersThis immediately tells you how many identity states may need to be controlled.
ExampleOwner, assistant, Nicola = three meaningful subjects.
Common mistakeDiscovering halfway through the storyline that another recurring subject needs a Character Sheet.
Approval testEvery recurring or narratively important person has been identified.
Define each subject's role.
A subject is not just a face. Define why they exist in the story.
Why it mattersRole determines whether they need continuity and how much visual attention they receive.
ExampleThe owner carries the problem; the assistant triggers the content request; Nicola introduces the system.
Common mistakeTreating every human figure as equally important.
Approval testYou can explain each subject's narrative job in one sentence.
Define visual states.
For each important subject, list the outfits or states that materially change.
Why it mattersDifferent outfits can require different Character Sheets even when identity is the same.
ExampleOwner in pajamas, then owner in suit.
Common mistakeUsing one Character Sheet across visually incompatible scenes.
Approval testEvery scene points to the correct visual state.
5.4 Write the scene as observable actions
A practical way to write is to think in verbs. What can the viewer literally see the person doing?
These verbs force clarity. They also make it easier to detect when a storyline contains an idea that has not yet been translated into visible action.
Subject Architecture and Character Logic
Four kinds of subject. Only the ones the audience must recognize get a Character Sheet.
Main subject
The person the audience is supposed to follow. Identity consistency is normally critical.
Why it mattersThis person should usually have a Character Sheet because visual drift damages trust and continuity.
ExampleBusiness owner appearing throughout the video.
Secondary controlled subject
A recurring person who interacts with the main subject and matters to the narrative.
Why it mattersIf this person appears repeatedly, speaks or performs a specific role, control them too.
ExampleAssistant who appears in multiple scenes.
Incidental subject
A person who exists only to make the scene feel natural and whose exact appearance does not matter.
Why it mattersLeaving these people uncontrolled keeps the reference stack simpler.
ExampleSleeping wife whose face is not shown.
Same identity, new visual state
The same person can require a new Character Sheet when outfit or visual state changes materially.
Why it mattersSeedance needs a clean visual anchor for the state used in that scene.
ExampleOwner in pajamas vs. owner in suit.
Automatically creating a Character Sheet for every human in the frame.
Ask whether the audience must recognize and track this person. If yes, control them.
Motion, Attitude and Performance Direction
Motion is the skeleton. Attitude is the performance. Write both.
Motion tells the model what happens. Attitude tells the model how it happens. Both are necessary for a convincing performance.
Motion
Describe the physical sequence: stand, reach, walk, hand over, refuse, turn, sit, drive, pick up, drop.
Why it mattersPhysical actions are the skeleton of video generation.
ExampleHe reaches toward the alarm, misses, knocks over the water, reaches again and switches it off.
Attitude
Describe the emotional quality of the movement: exhausted, annoyed, relaxed, confident, nervous, excited.
Why it mattersWithout attitude, correct actions can still feel dramatically wrong.
ExampleThe owner moves slowly and irritably after working for twelve hours.
Intensity
State whether the performance is subtle or exaggerated when this affects the result.
Why it mattersAI can over-act if the emotional direction is too broad.
Example“Visibly tired but realistic, not theatrical.”
Interaction
When two subjects interact, state who initiates the action and how the other responds.
Why it mattersMulti-subject scenes often fail when interaction order is ambiguous.
ExampleThe assistant extends the phone first; the owner looks at it, refuses, and pushes it back.
Using emotion words without translating them into visible behavior.
A viewer should be able to infer the emotion from the generated performance even with the sound off.
Location Logic and Spatial Consistency
One frame, a full sheet, or nothing at all — proportional to how much the room must be recognized.
Location is more than a background. It determines what the camera can plausibly see, where subjects can move and whether multiple shots feel like they belong to the same world.
Single-frame location reference
Use one reference when you only care about one specific view or setup.
Why it mattersIt is faster and sufficient when the camera does not need to understand a full room.
ExampleA podcast desk seen from one main angle.
Common mistakeBuilding a full Location Sheet when one strong reference already solves the scene.
Location Sheet
Use a fuller sheet when the camera will explore multiple angles or the architecture must remain coherent.
Why it mattersMultiple views help communicate the space beyond one flat image.
ExampleBedroom shown from bed, doorway and side angle.
Common mistakeExpecting one narrow photo to define an entire room.
Generic location
If the exact room is not important, describe it and let the model create it.
Why it mattersThis reduces asset creation and keeps production focused on what matters.
ExampleGeneric modern office used for one short shot.
Common mistakeReferencing a location purely out of habit.
The location choice should be proportional to how much the audience needs to recognize that exact space.
Object Logic and Prop Consistency
Reference an object only when it carries story, brand or continuity.
An object deserves a reference when it carries story, branding or continuity. Objects that are merely decorative can stay generic.
Narrative object
The object causes or advances an action.
Why it mattersIf it changes, the action can become confusing.
ExampleAlarm clock used in the wake-up hook.
Recurring object
The same prop appears across multiple shots.
Why it mattersConsistency helps the audience perceive one continuous story.
ExampleDossier handed over in one shot and held later.
Hero product
A product is part of the message or offer.
Why it mattersVisual accuracy can be strategically important.
ExampleA product package or branded device.
Generic prop
An object fills the environment but does not matter.
Why it mattersNo reason to spend reference budget on it.
ExampleA generic glass of water beside the alarm.
Creating a sheet for every prop visible in the frame.
If the object changed slightly and nobody would care, do not control it.
Camera Direction, Framing and Transitions
Static, zoom, shake, hard cut, tracking, angle change — used intentionally, never decoratively.
The recordings emphasize that the model can now create camera-angle changes and movement that previously required separate still images. Your job is to describe camera behavior clearly when it is part of the creative intention.
Static camera
Use when performance and dialogue matter more than movement.
Why it mattersA stable shot reduces unnecessary variation.
Sharp zoom
Use to direct attention quickly to a key action.
Why it mattersThe alarm sequence uses a sharp zoom to emphasize the reach.
Slight camera shake
Use sparingly to add physical energy or realism.
Why it mattersIt can make a frantic or reactive moment feel less artificial.
Hard cut
Use when the storyline jumps directly to a new action or time/location.
Why it mattersBedroom → shower → brushing teeth can be described as hard cuts.
Tracking / following
Use when the camera should maintain relation to a moving subject.
Why it mattersUseful for walking or moving through an environment.
Angle change
Use when the same action benefits from a different perspective.
Why it mattersThe new system can shift angle inside the generated sequence if the direction is clear.
Adding camera movement because it sounds cinematic rather than because it improves the scene.
The camera should serve the action and story.
References
The visual memory of the system: how to build, approve, number and upload Character, Location and Object references.
Reference Architecture: Character, Location and Object Sheets
Stable visual ingredients instead of final frames — and the one question that decides what gets referenced.
The reference architecture is the visual memory of the system. Instead of manufacturing the final frame, you manufacture stable visual ingredients.
| Element | Explanation |
|---|---|
| Character Sheet | Controls a person's identity and appearance across angles and actions. |
| Location Reference / Sheet | Controls where the scene happens and, when needed, how the space is structured. |
| Object Reference / Sheet | Controls a prop or object that must stay recognizable. |
11.1 The minimum-reference principle
Use the smallest set of references that still gives you the control you need. Too few references create drift. Too many references create unnecessary complexity.
Would the final video be meaningfully worse if the AI changed this element? If yes, reference it. If no, describe it.
Building and Approving Character Sheets
Six steps from source image to locked canonical state, plus the approval checklist.
Choose the source image carefully.
Start with a clear image where identity is readable. A weak source creates ambiguity before generation even begins.
Why it mattersThe Character Sheet can only be as reliable as the identity information supplied to it.
Use the standard Character Sheet prompt.
The workflow relies on a repeatable master prompt rather than improvising a new identity-sheet prompt every time.
Why it mattersStandardization makes team output easier to review and reduces operator-to-operator variance.
Generate in GPT Image 2.5 at 16:9.
The training specifies this setup for the Character Sheet workflow.
Why it mattersKeeping the format consistent simplifies reuse throughout the production system.
Add an outfit reference when required.
If you need a subject in a specific outfit, include that outfit as an additional reference and explicitly instruct the image model to dress the subject accordingly.
Why it mattersOutfit accuracy is easier to control during sheet creation than to hope for later in video generation.
Review the sheet before approving it.
Check face structure, skin tone, hair, facial hair, age, body proportions, outfit, accessories and any recognizable details.
Why it mattersA bad Character Sheet contaminates every later generation that uses it.
Lock the approved state.
Once a sheet is approved, treat it as the canonical visual state for that scene or outfit.
Why it mattersChanging references casually during production can cause the same person to drift across generations.
Approving a sheet because it looks attractive even though it does not accurately match the subject.
The sheet should maximize identity consistency, not artistic reinterpretation.
12.1 Character Sheet approval checklist
- Face is recognizably the same person.
- Hairline, hairstyle and hair texture are correct.
- Facial hair is correct.
- Age is stable and believable.
- Body proportions do not noticeably drift.
- Outfit matches the intended scene state.
- Accessories are consistent.
- Nothing important is added or removed.
- The sheet is clean enough to serve as a reliable video reference.
Building and Approving Location References
Start from what the camera needs, not from what looks nice.
Start from the shot requirement
Decide what the camera needs to understand before creating a location asset.
Why it mattersYou avoid over-building a location that will only ever be seen from one angle.
ExampleOne front-facing office frame may require only one image.
Use an existing visual when sufficient
A generated location is not mandatory. A suitable existing reference can be enough.
Why it mattersThe goal is to communicate visual intent, not to generate assets for the sake of generating them.
ExampleA Pinterest bedroom reference can define the space.
Create a Location Sheet when multiple angles matter
A multi-view location asset becomes valuable when the room itself must remain stable across several shots.
Why it mattersThis gives the model more spatial information.
ExampleA bedroom seen from bed, side wall and doorway.
Review spatial logic
Check whether the room makes sense as one place.
Why it mattersInconsistent geometry can create impossible transitions.
ExampleDoor, bed, window and furniture should not contradict one another.
Selecting a beautiful location reference that cannot support the camera moves in the storyline.
The reference should enable the planned scene, not merely look good.
Building and Approving Object References
One canonical object reference, reused every time the object returns.
Identify the hero object
Choose objects that the audience needs to recognize.
Why it mattersThese are the objects worth controlling.
ExampleThe dossier handed by the assistant.
Capture defining shape and materials
The reference should make the object visually unambiguous.
Why it mattersAI can mutate poorly defined props.
ExampleA dossier with clear cover, dimensions and color.
Avoid over-referencing generic items
Generic objects can be generated naturally.
Why it mattersThis keeps the asset set lean.
ExampleA normal glass of water.
Keep the same asset across generations
If the object returns, reuse the same reference.
Why it mattersThis prevents the prop from changing between clips.
ExampleSame dossier across multiple shots.
Creating a new version of the same hero object for every generation.
One canonical approved object reference should normally carry through the sequence.
Reference Numbering, File Naming and Upload Discipline
Storyline, Claude and Seedance must agree on what Reference 1 means.
Reference numbering is operational infrastructure. The storyline, Claude and Seedance must all agree on what Reference 1, Reference 2 and Reference 3 mean.
| Reference | Canonical meaning |
|---|---|
| Reference 1 | Subject 1 — owner in pajamas |
| Reference 2 | Bedroom / bed |
| Reference 3 | Alarm clock |
| Reference 4 | Subject 1 — owner in suit |
| Reference 5 | Assistant |
| Reference 6 | Office |
15.1 Recommended filename pattern
01_SUBJECT_OWNER_PAJAMAS.png 02_LOCATION_BEDROOM.png 03_OBJECT_ALARM_CLOCK.png 04_SUBJECT_OWNER_SUIT.png 05_SUBJECT_ASSISTANT.png 06_LOCATION_OFFICE.png
Number the files
Put the reference number at the beginning of the filename.
Why it mattersThe operating system will naturally sort them in the intended order.
Name the asset type
Use SUBJECT, LOCATION or OBJECT.
Why it mattersOperators can identify the reference at a glance.
Name the state
Add outfit or scene state when necessary.
Why it mattersThis avoids using the correct person in the wrong visual state.
Do a pre-upload map check
Read the storyline and compare every reference callout to the actual file list.
Why it mattersThis catches the most preventable generation error before credits or time are spent.
Relying on visual memory when several similar files are open.
Someone who did not create the files should still be able to upload them correctly.
Claude and Seedance 2.5
Translating the storyline into a prompt, splitting it into generations, and running them.
Claude as the Translation Layer
Claude encodes your decisions. It does not make them.
Claude is used after the human creative system is defined. It receives the storyline plus the numbered visual references and converts them into a Seedance-focused prompt.
Paste the storyline segment, not random notes.
Claude should receive one coherent creative instruction set.
Why it mattersClean input produces cleaner structured output.
Upload the references in the exact numbered order.
Claude needs to understand which visual corresponds to each mention in the storyline.
Why it mattersThis preserves the reference map.
Specify generation duration.
The prompt has to fit the amount of time available.
Why it mattersTen seconds and thirty seconds require different pacing.
Specify editing energy when useful.
High-energy and polished-cinematic are different execution styles.
Why it mattersThis helps Claude shape cuts, pacing and camera behavior.
Review Claude's output.
Claude handles technical detail, but you still own the creative intent.
Why it mattersCheck references, action order, timing and camera logic before pasting into Seedance.
Treating Claude output as automatically correct because it sounds technically sophisticated.
The prompt is approved only when it faithfully encodes the storyline.
16.1 What Claude should preserve
- Subject identity and numbering.
- Location identity and numbering.
- Object identity and numbering.
- Action order.
- Subject attitude and emotion.
- Camera movement.
- Transitions and hard cuts.
- Generation duration.
- Overall editing energy.
Splitting a Video into Seedance Generations
Two methods, one rule: split where the story naturally breaks.
The training workflow uses Seedance 2.5 with generations up to 30 seconds. A longer video therefore needs multiple generations, and some shorter sequences should still be split for control.
Method A — Let Claude split the full storyline
Give Claude the complete sequence and ask for a specific number of generation prompts.
Why it mattersFast when timing does not need to be micromanaged.
ExampleSplit this into three Seedance generations.
Method B — Split manually first
Decide exact segments before prompting Claude.
Why it mattersBest when dialogue, beats or scene transitions need precise timing.
ExampleGeneration 1: 10s wake-up; Generation 2: 5s bathroom; Generation 3: 20s office.
Split at natural story boundaries
A generation should ideally contain a coherent unit of action.
Why it mattersNatural boundaries reduce continuity problems and make review easier.
ExampleEnd one generation after the wake-up sequence, begin the next at the shower hard cut.
Avoid forcing too much into one generation
More actions are not automatically better.
Why it mattersOverloaded prompts can produce rushed or skipped beats.
ExampleIf one 30-second block contains too many location changes, split it.
Choosing splits only to minimize the number of generations.
Choose splits that maximize narrative control and generation reliability.
Seedance 2.5 Execution Workflow
Eight steps, in order. When a clip fails, change one thing.
- Step 1: Open Seedance 2.5 in Higgsfield/Xfield. — Use the model selected in the workflow.
- Step 2: Paste the complete Claude prompt. — Do not remove reference logic unless you intentionally change the scene.
- Step 3: Upload the required references. — Only upload the assets used by that generation.
- Step 4: Verify the mapping. — Image 1 must be what the prompt calls Image/Reference 1.
- Step 5: Set the intended duration. — Match the duration used when Claude structured the prompt.
- Step 6: Generate. — Run the clip.
- Step 7: Review against the storyline. — Check whether the story was executed, not just whether the clip looks high quality.
- Step 8: Diagnose before rewriting. — If something failed, identify whether the problem was identity, reference mapping, action, timing, camera or storyline clarity.
Each step protects a different part of the pipeline.
Changing several variables after one bad generation.
Change the smallest thing that plausibly caused the failure.
The operating procedure
The full SOP from concept to assembly, the review framework, and what to do when something breaks.
Full End-to-End SOP
Concept → Storyline → References → Claude → Seedance → Assembly. Every item is a decision.
Every item below is a required part of its phase. Treat each one as a production decision, not a box to tick. The purpose is to reduce ambiguity before the next phase begins.
Phase 1 — Concept
- Write the topic and business purpose of the video.
- Write the idea in one or two sentences.
- Write the script or dialogue.
- Decide the rough format: simple talking video, ad, podcast-style clip, multi-scene story, or other.
Phase 2 — Storyline
- List the subjects.
- Define each subject's role.
- Define visual states/outfits.
- Describe physical actions.
- Describe attitude/emotion.
- Choose locations.
- Identify important objects.
- Describe camera behavior.
- Write the scene progression from first frame to final beat.
- Add explicit cuts and transitions.
Phase 3 — References
- Create approved Character Sheets.
- Create additional Character Sheets for changed outfits.
- Create or source location references.
- Create Location Sheets only when needed.
- Create object references for hero props.
- Assign reference numbers.
- Rename and organize files.
Phase 4 — Claude
- Open the Seedance Skill.
- Paste the relevant storyline.
- Upload references in exact order.
- Set duration and editing energy.
- Generate the Seedance prompt.
- Review and approve the prompt.
Phase 5 — Seedance
- Paste the Claude prompt.
- Upload references in matching order.
- Verify mapping.
- Generate.
- Review against storyline.
- Adjust the specific failure source if needed.
Phase 6 — Assembly
- Generate remaining sections.
- Place clips in sequence.
- Review continuity between clips.
- Add voice, audio, captions or sound design.
- Run final quality control.
Moving forward while an item is still undecided.
The next operator should be able to continue without asking what you meant.
Quality Control and Review Framework
Eight things to check on every generation before you approve it.
Identity consistency
Does the subject still look like the approved Character Sheet?
Why it mattersCheck face, hair, age, body proportions, outfit and accessories.
Location consistency
Does the environment still match the intended location?
Why it mattersCheck architecture, furniture, light logic and spatial coherence.
Object consistency
Do hero objects keep the same identity?
Why it mattersCheck shape, size, color and material.
Action accuracy
Did the requested actions actually happen?
Why it mattersCheck order, interactions and whether any beat was skipped.
Attitude accuracy
Does the performance feel like the intended emotion?
Why it mattersCheck body language, pace and expression.
Camera accuracy
Did the camera behave in a way that supports the scene?
Why it mattersCheck framing, movement, angle and cuts.
Narrative accuracy
Does the clip communicate the intended story beat?
Why it mattersThe best-looking clip can still be wrong if it tells a different story.
Continuity
Does this generation connect logically to the one before and after it?
Why it mattersCheck wardrobe, object state, subject position and story progression.
Approving based on visual quality alone.
Approve only when both visual quality and narrative accuracy are acceptable.
Troubleshooting and Failure Modes
Symptom, diagnosis, fix. Change the layer that broke, not the whole system.
The face drifts.
DiagnosisLikely causes include a weak Character Sheet, insufficient identity information, wrong reference mapping or too much visual conflict.
FixRecheck the canonical Character Sheet and the upload order before changing the storyline.
The outfit is wrong.
DiagnosisThe wrong Character Sheet state may have been used or the outfit was not locked during sheet creation.
FixUse a dedicated Character Sheet for the intended outfit.
The location changes too much.
DiagnosisOne reference may not be enough for the camera coverage requested.
FixUse a clearer or fuller Location Sheet if spatial continuity matters.
The object morphs.
DiagnosisThe prop may be under-specified or not referenced.
FixCreate a cleaner canonical object reference.
The model skips an action.
DiagnosisThe generation may be overloaded or the action sequence may be vague.
FixSimplify the generation or rewrite the action as explicit steps.
The camera behaves randomly.
DiagnosisThe storyline may not contain clear camera direction.
FixSpecify static, zoom, hard cut, tracking or angle changes where they matter.
The clip looks great but feels wrong.
DiagnosisThe aesthetic quality is high but the narrative beat is off.
FixReturn to the storyline and compare what was requested with what occurred.
Different generations do not connect.
DiagnosisThe boundary between segments may be poorly chosen or states may not match.
FixSplit at natural transitions and carry the same canonical references across segments.
Claude produces an overcomplicated prompt.
DiagnosisThe input may contain too many unresolved decisions.
FixSimplify the storyline and provide cleaner constraints before prompting again.
Production feels slow again.
DiagnosisThe team may be rebuilding the old image-first workflow inside the new system.
FixAsk whether every asset being generated is truly required as a reference.
Example, templates, checklists
A complete worked example, the three copy-paste templates, the final checklists and the operating rules.
Complete Worked Example: Busy Business Owner
The whole system applied to one video: three subjects, two outfits, three locations, hard cuts and a hero prop.
22.1 Core concept
A busy professional has been working since early morning and has no time or energy to record content. His assistant asks him to film an Instagram video. He refuses. The story then introduces the AI-clone solution and shows that the team can create content without requiring him to record.
22.2 Subject map
| Subject | Function |
|---|---|
| Subject 1 | Business owner — main character and problem carrier. |
| Subject 2 | Assistant — triggers the request to record and represents the team. |
| Subject 3 | Nicola — introduces or teaches the AI-clone solution. |
| Incidental wife | Background contextual subject only. Exact identity is not important. |
22.3 Initial reference map
| Reference | Asset |
|---|---|
| Reference 1 | Owner Character Sheet — pajamas. |
| Reference 2 | Bedroom / bed location. |
| Reference 3 | Alarm clock. |
| Reference 4 | Owner Character Sheet — work outfit/suit. |
| Reference 5 | Assistant Character Sheet. |
| Reference 6 | Office location. |
22.4 Expanded opening storyline
The video opens before the workday begins. Subject 1 (Reference 1) is asleep in the bedroom bed (Reference 2). The framing should immediately establish fatigue rather than luxury or relaxation. Beside him is his wife, a blonde woman who is not a controlled subject. Her exact face is not important because the audience should remain focused on Subject 1.
Suddenly, the alarm clock (Reference 3) on the nightstand goes off. The camera makes a sharp zoom toward the side of the bed with a small amount of camera shake. Subject 1 reaches toward the alarm without fully waking up. On the first attempt, his hand misses the alarm and knocks over a glass of water. The mistake should feel natural, not comedic or exaggerated. He reaches again, grabs the alarm and switches it off. His body language communicates that he is exhausted and already irritated by the day.
Hard cut: Subject 1 is in the shower. Hard cut: he brushes his teeth in front of the mirror. Hard cut: he prepares for work. When the outfit changes materially, the generation should use the Character Sheet created for the work outfit rather than the pajamas sheet.
The story then transitions to the office. The assistant approaches and asks him to record content. The assistant initiates the interaction by presenting the phone or content request. Subject 1 looks at it, immediately refuses, and returns attention to work. His refusal should feel like the reaction of someone who genuinely has no time left, not like a staged advertisement.
The AI-clone solution is then introduced. At this point the story changes from problem to mechanism: the team can create the content while the owner remains focused on the business.
This example demonstrates the full logic of the system. The video contains several locations, multiple subjects, different outfits, camera changes, hard cuts and important props, yet the creator does not need to generate every final frame manually. The system is built from a storyline plus a controlled reference stack.
Reusable Templates and Production Forms
Three templates. Copy them, fill the brackets, keep the structure.
VIDEO IDEA [One-sentence concept] SCRIPT / DIALOGUE [Spoken content] SUBJECT MAP Subject 1 — [role] Subject 2 — [role] Subject 3 — [role] VISUAL STATES Subject 1 State A — [outfit/state] Subject 1 State B — [outfit/state] REFERENCE MAP Reference 1 — [asset] Reference 2 — [asset] Reference 3 — [asset] STORYLINE The video opens with... Then... Camera... Subject attitude... Hard cut... Next action... End of generation... GENERATION PLAN Generation 1 — [duration / content] Generation 2 — [duration / content] Generation 3 — [duration / content] EDITING ENERGY [High energy / polished cinematic / other]
1. Is this element important to the story? 2. Will the audience notice if it changes? 3. Does it repeat across shots? 4. Does exact identity matter? 5. Does exact outfit/state matter? 6. Does exact location geometry matter? 7. Does the object carry narrative or brand meaning? If the answer is YES to one of the important consistency questions, create or use a reference. If the answer is NO, describe the element and let the model create it.
I want to create the following video in Seedance 2.5. Use the storyline below and the uploaded references in the exact numbered order. Convert it into a production-ready Seedance prompt. Preserve subject identities, locations, objects, action order, attitude, camera direction, scene progression and timing. STORYLINE: [PASTE] GENERATION DURATION: [INSERT] EDITING ENERGY: [INSERT]
Final Checklists
Before references, before Claude, before Seedance, before final approval.
24.1 Before creating references
- The idea is clear.
- The script is stable enough to plan.
- The storyline exists.
- The subjects are identified.
- The visual states/outfits are identified.
- The locations are identified.
- The hero objects are identified.
- The camera behavior is described where it matters.
24.2 Before Claude
- Every controlled subject has the correct Character Sheet.
- Location references are approved.
- Object references are approved.
- Reference numbers are final.
- File names match the reference numbers.
- The relevant storyline section is final.
- The generation duration is known.
- The editing-energy direction is known when needed.
24.3 Before Seedance
- Claude prompt has been reviewed.
- Reference order is correct.
- Prompt numbering matches upload numbering.
- Duration matches the generation plan.
- Only the references needed for this generation are being used.
- The intended beginning and end of the clip are clear.
24.4 Before final approval
- Identity is consistent.
- Outfit state is correct.
- Location is consistent.
- Hero objects are consistent.
- Actions occurred in the right order.
- Emotion/attitude feels right.
- Camera behavior supports the scene.
- The clip advances the intended story.
- Transitions into adjacent generations work.
- Visual quality is high enough for final use.
Operating Rules and Glossary
Fourteen rules that summarize the whole system, and the vocabulary the team shares.
14 operating rules
- Storyline first. Do not generate assets for a scene you have not defined.
- Control important visual DNA; delegate unimportant detail.
- One important subject normally equals one Character Sheet per meaningful outfit/state.
- Use a single location image when it is enough; use a fuller Location Sheet when spatial continuity matters.
- Reference a prop only when its consistency matters.
- Keep reference numbering identical across storyline, Claude and Seedance.
- Describe observable actions rather than abstract ideas.
- Describe attitude as well as movement.
- Use camera language intentionally, not decoratively.
- Split videos at natural story boundaries.
- Claude translates. You direct.
- Seedance generates. You judge.
- Review against the storyline, not against visual beauty alone.
- When something fails, diagnose the exact layer before changing the whole system.
Glossary
| Term | Meaning |
|---|---|
| Storyline | The detailed visual sequence of the video. |
| Subject | A person or character appearing in the scene. |
| Visual state | A version of a subject defined by outfit or meaningful appearance state. |
| Character Sheet | The canonical multi-view visual identity reference for a subject/state. |
| Location Reference | An image that defines the environment. |
| Location Sheet | A broader multi-view representation of a location. |
| Object Reference | An image that defines a hero prop. |
| Object Sheet | A broader visual reference for an object requiring strong consistency. |
| Reference number | The number linking the same asset across storyline, Claude and Seedance. |
| Generation | One generated video segment. |
| Editing energy | The pacing and visual execution style requested for the segment. |
| Hard cut | An immediate jump from one shot/action to another. |
| Canonical asset | The approved version of a reference that the team reuses consistently. |
Each layer feeds the next. Keep them separate, finish each one before moving on, and the system stays fast.
Build the story. Lock what matters. Let the model create the scene.
- Storyline first — always
- Control what matters, delegate the rest
- Same reference numbers everywhere
- Review against the storyline, not beauty
IDEA → SCRIPT → STORYLINE → REFERENCES → CLAUDE → SEEDANCE 2.5 → REVIEW → FINAL VIDEO
Want to skip the learning curve and learn directly from me?
I'll build your AI clone and this exact storyline-first system with you — so your team produces the videos and you never record again.
Apply to work with me → Instagram — @nicola.ai