---
title: Clone Yourself
description: What to record and photograph so Wireflow can make an AI twin of you that sounds like you, looks like you and moves its hands like you. Voice, face, mouth, a talking clip and a gesture clip, in that order, and how each one feeds the tested avatar recipe.
updated: 2026-09-29
---

An AI twin is only as good as what you give it. Everything below is stuff you record once, on your phone, in about 20 minutes. After that every new video reuses it.

Do the captures in this order. Each one fixes a specific thing that goes wrong without it, and the reasons are listed so you know what to redo if a result looks off.

Using an agent (Claude Code, Claude, ChatGPT)? It can read this same checklist with MCP `get_recipe` and slug `built-in-digital-twin-setup`. One catch: chat clients like claude.ai and ChatGPT can't hand your files to Wireflow yet. Upload them in the web app with an Import node, or use Claude Code, which can upload through the API.

## 1. Voice

This is what the clone learns your voice from. Wireflow sends it to ElevenLabs Instant Voice Cloning.

What ElevenLabs says to record (checked 2026-09-29 against their [Instant Voice Cloning guide](https://elevenlabs.io/docs/product-guides/voices/voice-cloning/instant-voice-cloning), their [voice cloning overview](https://elevenlabs.io/docs/product-guides/voices/voice-cloning) and their [Professional Voice Cloning guide](https://elevenlabs.io/docs/product-guides/voices/voice-cloning/professional-voice-cloning)):

- **1 to 2 minutes of speech.** More than 3 minutes barely helps and can make it worse. Wireflow's node takes 30 to 120 seconds.
- **A quiet room, no echo.** No background noise and no reverb, they end up in the clone. If the room echoes, a thick duvet or quilt around you helps.
- **Same distance from the mic the whole time.** About two fists away. Don't lean in and out, and keep your volume steady.
- **Talk the way you want the clone to sound.** It copies your delivery, so read naturally, like you're explaining something to a friend. Keep it steady though: big swings in pitch or volume, long silences and shouting all come through.
- **Just you.** No second voice, no music, no sound effects, no processing on the file.
- **File:** MP3 at 128 kbps or better (192 kbps if your recorder offers it). Wireflow's node also takes M4A, up to 25 MB.

Then in Wireflow: add an Import node, upload the file (or hit **Record audio** in the node), and wire it into **Voice Sample** on an [ElevenLabs Voice Clone](/docs/nodes/audio--elevenlabs_voice_clone) node (`audio:elevenlabs_voice_clone`). Name the voice, tick the consent box, run it once. It costs 50 credits and gives you a `voice_id` you reuse forever.

To talk in it, wire that `voice_id` into [ElevenLabs TTS (Custom Voice)](/docs/nodes/audio--elevenlabs_tts_custom_voice) (`audio:elevenlabs_tts_custom_voice`) with Model set to `eleven_v4`. Write the script the way people talk ("it's", "didn't"), not "it is". If it sounds a bit bright or harsh, run it through [Audio EQ](/docs/nodes/utility--audio_eq) with the `voice_deharsh` preset.

## 2. Face

Real photos are the best identity refs there are. Take 4 to 6:

- **Front on and a three-quarter turn**, at least one of each. Head and shoulders, face big in the frame.
- **Good soft light from the front.** A window in front of you is ideal. No hard shadows across the face.
- **No filters, no beauty mode, no portrait blur.** Turn them off in the camera app.
- **Relaxed face, mouth closed or barely open.** No big posed smile.
- **Look like you normally look** in videos: same glasses, same hair, same beard.

## 3. Mouth

Two close-ups of your real mouth, cropped tight on lips and teeth, same light as the face photos:

- one with your lips relaxed and closed
- one mid-word, teeth showing the way they naturally do when you talk

Without these the video model invents a mouth, and it tends to show too much upper teeth and gum. This is the single most common "that's not me" tell. Both go into the take.

## 4. Talking clip

A selfie video of you talking normally, phone at eye level, arm's length, light in front of you. This teaches the model how your mouth, teeth and jaw actually move. It is a reference only and never ends up on screen.

- Record 10 to 20 seconds so you have a good bit to pick from, then cut a **3 to 4 second** window of you mid-sentence. That short window is what goes into the take (the tested take used 3 s).
- Talk like you do on camera. Not performing, not reading slowly.
- Face straight at the lens most of the time.
- Record it in the shape you want your videos (portrait for reels). The video model follows the reference clip's orientation.

## 5. Gesture clip

This is what makes the hands move. Without it the twin keeps whatever hand pose the still has, frozen for the whole clip. With it, the model copies the timing of your own hand and torso movement.

- **Sit at a desk.** Webcam, or your phone propped up on the desk.
- **Head to desk in frame.** Your hands have to be visible when they rest on the desk.
- **Light from the front.**
- **Rest, lift, rest.** Hands resting apart on the desk, then a natural gesture while you talk, then back to rest. Mostly one hand.
- **Slow.** Fast moves blur, and blur teaches nothing.
- **Quiet room.** Audio in a reference clip can leak into the voice of the take. The tested gesture clip had no sound at all. Wireflow can't strip the sound off a clip yet (tracked: [#2311](https://github.com/wireflowINC/wireflow/issues/2311)), so record it somewhere quiet.
- Record 15 to 20 seconds, then cut an **8 second** window with the gesture where you want it. The hands move when the clip's hands move; Seedance ignores "lift the hand on this word" in the prompt.

Start with your hands resting apart, not clasped. The first frame of the video tends to match the still's hand pose, and the gesture clip only frees the hands if it shows them open.

## 6. Optional: borrowed gestures

No gesture clip of your own? You can point at a clip of someone else talking with their hands. It does make the hands move. The honest catch: borrowed gestures can bring their stuff along. In our test another creator's ring showed up on the twin's hand. Your own clip is best, and it's 20 seconds of effort.

## How it all plugs into the recipe

The tested method is in [Make a Realistic AI Avatar](/docs/make-a-realistic-ai-avatar). Here's where each capture goes. This is the exact setup of the take we locked (called take D below).

**The still.** One [GPT Image 2.5 Sunburst](/docs/nodes/generate--openai_gpt_image_2_5_sunburst_text_to_image) pass at custom 1152x2048, built from your face photos as identity refs, with the scene described in the prompt. A face-free scene plate goes first on `image1`, your photos after it. No upscale, no grain words. Ask for hands resting apart on the desk.

**The voice line.** Write the line, speak it with your cloned voice (ElevenLabs TTS, `eleven_v4`), then run it through [Audio Pad](/docs/nodes/utility--audio_pad) (`utility:audio_pad`) with 0.5 s of silence before and after, 48 kHz stereo. That padded file is `@Audio1`.

**The video.** [Seedance 2.5](/docs/nodes/video--seedance_2_5) in reference mode, Start Frame left empty, **Generate Audio** on. Wire the still first, check each `@ImageN` badge, then give every reference one plain job in the prompt:

| Reference | What it is                 | Where it's wired | Its job in the prompt                                               |
| --------- | -------------------------- | ---------------- | ------------------------------------------------------------------- |
| `@Image1` | the still you just made    | Reference Images | the shot: framing, lighting, hair, clothes, background              |
| `@Image2` | your front photo           | Reference Images | controls only the face, don't copy hair, clothes or lighting        |
| `@Image3` | your three-quarter photo   | Reference Images | controls only the face, don't copy hair, clothes or lighting        |
| `@Image4` | mouth close-up, lips shut  | Reference Images | mouth shape only: how the lips move and how little upper teeth show |
| `@Image5` | mouth close-up, mid-word   | Reference Images | mouth shape only, same as `@Image4`                                 |
| `@Video1` | your 3 to 4 s talking clip | Reference Video  | only head movement, mouth articulation and expression timing        |
| `@Video2` | your 8 s gesture clip      | Reference Video  | "the speaker's natural hand and torso gesture timing"               |
| `@Audio1` | your padded voice line     | Audio            | only the words, timing and rhythm of the speech                     |

Then write the script into the prompt as a **timeline**: each line with its time window ("0.7 to 1.7 s: he says ..."), and what the face does in every silent gap ("lips together, relaxed, no teeth showing"). Keep mouth wording on the mouth refs only. Keep "add", "remove", "replace", "extend" and "continue" out of the prompt.

**Seedance re-voices `@Audio1`.** It uses your line for the words and timing, but the voice that comes back is its own take on it, close to you but not your exact clone. For your exact voice, run [Sync Lipsync v3](/docs/nodes/talking--sync_lipsync_v3) (`talking:sync_lipsync_v3`) on the finished take with the voice line.

A few limits that bite:

- **The take is as long as your clips.** With a reference clip wired, Seedance ignores the Duration setting and the take comes out about as long as your longest clip. Take D used a 3 s talking clip and an 8 s gesture clip and came out 8 s. You pay for the take's seconds plus the reference footage seconds, so a 20 s clip means a roughly 20 s take and a bigger bill. Cut clips to the window you want with [Video Trim](/docs/nodes/utility--video_trim) (free).
- **Both clips together stay at or under 30 seconds.** That's Seedance's cap on reference footage.
- **This is 7 reference materials** (5 images and 2 clips; the voice line doesn't count). The normal run path takes up to 8. One older path takes only 5 and refuses the run before charging you. If you hit that, drop the three-quarter photo and one mouth close-up.

Take D ran at 480p. 720p is optional and costs more per second. Ship the raw take either way, no upscale. Grade and grain, if you want them, go on the final cut in the editor.

## Keep it private

A face and a voice identify you. Keep your captures and your cloned voice in your own account, and never publish a real person's face or voice as a public blueprint. Clone only your own voice, or one you have permission to use.

---

Documentation index: fetch https://www.wireflow.ai/llms.txt for the full list of Wireflow docs.
