---
title: Make a Realistic AI Avatar
description: The tested recipe for a photoreal AI talking-head avatar on Wireflow, a single GPT Image 2.5 Sunburst still animated with Seedance 2.5 in reference mode. Includes the mistakes that cost real credits, with the numbers.
updated: 2026-09-30
---

This is the method that actually holds up, measured across real production runs, not a first guess. It applies to a founder talking to camera or a client avatar, really any AI creator that has to survive someone looking closely. Follow it in order: still first, then motion. Skipping straight to video from a bad still just wastes Seedance credits on a face that was never going to hold.

## The still: one GPT Image 2.5 Sunburst pass

Use [GPT Image 2.5 Sunburst](/docs/nodes/generate--openai_gpt_image_2_5_sunburst_text_to_image) (`generate:openai_gpt_image_2_5_sunburst_text_to_image`). Wiring any image into it routes the run to [the edit variant](/docs/nodes/edit--openai_gpt_image_2_5_sunburst_edit) (`edit:openai_gpt_image_2_5_sunburst_edit`) automatically, with the first wired photo landing on `image1`. That single fact decides how you wire identity refs.

1. **Never let a face photo land on `image1`.** Whatever is wired first becomes the edit canvas, and the model inherits that photo's exact pose, light and framing, so every result looks like a variation of one frame. Two ways to avoid it:
   - No image wired at all: describe the whole shot (subject, pose, outfit, background, lighting) in the prompt and generate from text alone.
   - Identity refs needed: wire a face-free scene or style plate to `image1` first, then wire the identity photos into `image2`, `image3` and so on (each additional image you connect mints the next numbered slot). Label each identity ref in the prompt as "identity only, ignore pose, light, expression and clothing," and describe the actual shot in the prompt, not in any of the images.
2. **One face per pass.** If two people are in the shot, run two separate passes, each with only that person's refs wired, then composite. Loading both people's refs into one pass blends their features.
3. **Quality max, `image_size: custom` at 1152x2048, one pass.** That is the size that renders clean. The vendor's own reliable ceiling is 2560x1440; past that is experimental and shows up as over-sharpened skin or an etched beard, sometimes with a haloed edge around the hair.
4. **Do not upscale a still that feeds Seedance.** An upscale pass looks like more detail but it is more grain: a Crystal x1.5 pass measured 1.12 to 2.15 on a high-pass grain metric in testing, against 0.37 for a real iPhone photo. Seedance renders at 720p anyway, so a bigger still buys nothing downstream on that path; the pixels get thrown away. A still headed somewhere else, posted as an image in its own right, can still be upscaled the normal way (Crystal Upscaler, creativity 0).
5. **Keep "film grain," "sensor noise" and "pores" out of the prompt.** Every old plate prompt asked for grain in words, and paired with an upscale pass that itself added grain, the two together are what made past plates look synthetic. Dropping the wording is one half of the fix; skipping the upscale (rule 4) is the other.
6. **Keep the face large: waist-up or closer, never full body.** One test found a face about 40px tall in a 1024x1536 sheet lost identity; published benchmarks separately report similarity dropping from about 0.71 on a large face to about 0.56 on a small one. If just a couple of panels on a sheet come out wrong, fix only those with a masked edit (`edit:openai_gpt_image_2_5_sunburst_edit`, `mask_url` port) instead of re-rolling the whole sheet.

## The video: Seedance 2.5, reference mode, raw output

Once the still holds up, animate it with [Seedance 2.5](/docs/nodes/video--seedance_2_5) (`video:seedance_2_5`). "Reference mode" is not a setting on the node, it is a wiring pattern:

- **Leave Start Frame (`first_frame_url`) empty.** (`image_url` is a different port: a plain reference image, not a start frame.) Start Frame cannot be combined with reference media. With Reference Images or Reference Audio the run is refused before you are charged. With a Reference Video wired, the Start Frame is silently turned into one more reference image instead.
- **Wire the approved still, a front crop, a three-quarter crop, and a real mouth close-up into Reference Images (`reference_image_urls`).** Reference images and the Reference Video share one budget of registered materials, and some run paths cap it at 5. Four images plus the talking clip is exactly 5, which is why this version uses one mouth close-up. The set this recipe was tested with used two (6 materials), which those paths refuse before you are charged.
- **For a talking head, also wire a short real clip of the person talking into Reference Video (`video_url`) as `@Video1`.** Its pixels never end up on screen, but wiring a Reference Video makes Seedance run an edit task: the node's Duration and Aspect Ratio are ignored and the take follows the clip's length and orientation. Record the talking clip at the length and orientation you want the take to be. Without a real mouth close-up and a real talking clip, Seedance invents its own mouth shape and exaggerates the upper teeth and gums.
- **Wire the still first, and check the `@ImageN` badge on each port before writing the prompt.** The numbering follows wire order, not which slot you meant it for, so wiring out of sequence silently shifts every `@Image` reference after it.
- **Bind roles in the prompt with `@Image` and `@Video` references**: each scoped to its own job. `@Image1` is the composition, wardrobe and background reference. `@Image2` / `@Image3` are the identity crops, told to "control only the face, do not copy hair or clothing." `@Image4` is the mouth close-up, told to control "mouth appearance only." `@Video1` is the talking reference, told to control "timing and mouth movement only." Keep mouth wording on the mouth references only. A mouth instruction applied to every reference produced a fixed grimace, teeth bared even while silent.
- **Keep "add," "remove," "replace," "extend" and "continue" out of the prompt.** Seedance classifies its task from the prompt text, and any of those words, even inside a "do not copy hair" style instruction, can tip the run into an editing or extension task that then requires a different aspect-ratio setting and fails after submission if you did not set it.

This is what holds up through a head turn. First-frame mode (a still wired to Start Frame alone, no reference images) holds a frontal face fine but drifts identity as soon as the head turns.

Ship the raw 720p output. No video upscale by default: of the upscalers tested, some made it visibly worse (a waxy softened look, or an over-sharpened, processed look), and the one that came out closest to native did not do enough to justify the extra credits. The model that scored best on a sharpness or shimmer metric was picked worst in a human side-by-side in one such test. Sharper is not more real. Judge a finish by eye, never by a detail or shimmer score.

Grade and grain, if you want either, go on in the video editor afterward, applied to the whole cut, not baked into the plate or the raw clip.

## How to judge the result

Metrics screen out breakage, things like morphing teeth or drifting identity, plus visible shimmer. They do not tell you what looks real, and are not a way to pick a "winner."

- **Always compare side by side**, the candidate next to a real photo or a real clip of the same person, not the candidate alone.
- **Re-encode to phone quality before judging.** Instagram and TikTok re-encode everything; a clean master can hide grain or artifacts that only show up after a real upload. A candidate carrying 1 percent grain still shows about half of it after an IG re-encode, so judge the compressed version, not the raw export.
- **Trust your eye over the metric.** A sharper, more detailed frame reads as processed to a human, which is the opposite of the goal. If the metric's pick and your eye's pick disagree, go with your eye.
- **Check the face at crop zoom**, not just at the full frame. Small-frame likeness and crop-zoom likeness are different questions, and only the second one matters once the shot gets cut in tight.

## Common mistakes, with the numbers that proved them

| Mistake                                                                     | What happened                                                          | Measured                                                                                |
| --------------------------------------------------------------------------- | ---------------------------------------------------------------------- | --------------------------------------------------------------------------------------- |
| Upscaling a still that feeds Seedance                                       | Grain roughly doubled                                                  | 1.12 to 2.15 on a high-pass grain metric; a real photo measures 0.37                    |
| "Film grain" / "sensor noise" / "pores" in the prompt, plus an upscale pass | Old plates written and processed this way looked synthetic             | Dropping the wording and skipping the upscale, together, fixed it                       |
| Picking a finish by a sharpness or shimmer metric                           | In one test, the metric's "winner" was the worst pick by eye           | A human side by side overruled that metric's pick                                       |
| Face photo landing on `image1`                                              | The result looked like a derivative of that one photo's pose and light | Fixed by keeping identity refs off `image1`: a face-free plate there, refs in `image2+` |
| Full-body or wide framing                                                   | Identity fell apart once the face got small                            | About 40px tall in a 1024x1536 sheet lost identity in testing                           |
| Two people's identity refs in one pass                                      | Features blended between the two subjects                              | Fixed by one face, one pass, composite afterward                                        |

## When you need a Passport, and when real photos are enough

A [Passport](/docs/workflows-blueprints-playbooks#characters-a-face-and-voice-passport) is a saved identity reference board: a small set of approved photos (and, for a video model, front plus three-quarter crops) kept with the workflow so every new shot draws from the same refs instead of a fresh photoshoot each time.

Real people with real photos do not need one. Real photos are the best identity refs there are, so if you already have several real, varied shots of the subject, just wire those.

Invented characters need a Passport, because there is nothing else to hold identity across shots: generate a reference sheet once, lock it, and reuse it for every future still and clip. A Passport built by GPT Image 2.5 is fine to reuse this way, since grain did not climb pass over pass in testing. Identity does drift a little each time you generate a new sheet from an old one though, so keep drawing from the most recently approved sheet rather than a chain of regenerated ones.

A Passport identifies a real person by face and voice. Keep it private: fork it as a workflow rather than publishing it as a blueprint, and never publish a real person's face or voice as a public blueprint.

---

Documentation index: fetch https://www.wireflow.ai/llms.txt for the full list of Wireflow docs.
