I compared HeyGen's Photo Avatar, Digital Twin, Avatar V, and Instant Avatar on the same scripts. Avatar V held likeness best. See tests, costs, and picks.
A real estate agent I coach asked me a simple question last month: should she build a Photo Avatar, a Digital Twin, an Avatar V avatar, or an Instant Avatar? She had seen all four names in HeyGen tutorials and assumed they were four competing products.
They overlap more than the names suggest. Photo Avatar and Digital Twin are two ways to create an avatar of yourself.
Avatar V is the model that renders a Digital Twin. Instant Avatar is the earlier name for the video-avatar flow that became the Digital Twin.
So I ran the comparison she wanted.
Over two weeks, I built all four from the same person, then fed each one the same market-update script, a 10-minute training module, and a translation test. This guide covers what each avatar type is, how they compare feature by feature, three head-to-head tests, credit costs, and which one fits your situation.
Quick Verdict
Winner: a Digital Twin rendered on Avatar V. It held her face, teeth, and gestures steady across every test.
A Digital Twin on Avatar IV also learns real gestures but drifts over long scripts. Photo Avatar starts from one photo but invents its motion. Instant Avatar is the legacy flow that now runs as a Digital Twin.
Feature Comparison at a Glance
Avatar V
I recorded the agent once: 15 seconds on a laptop webcam, a live consent clip, and a two-minute expressive take as the motion reference. Avatar V became the default render engine the moment her Digital Twin finished training.
Her habit of tilting her head before a key point carried into the first render, and her teeth looked identical at second 10 and minute 9. That detail matters because general video models redraw a face slightly differently every few seconds.
Avatar V treats identity preservation as the design constraint, and it trains voice together with facial motion so mouth shapes follow the audio. It weighs audio first, the input image's expression second, and the text prompt third, which is why it suits an ai video avatar built from expressive footage.
Key Strengths
- Identity Lock: Avatar V scored 0.840 face similarity against Veo 3.1's 0.714, and human raters preferred it in 68.9% to 85.7% of head-to-heads against Kling, Seedance, Veo, and OmniHuman.
- Gaze Control: The model keeps the gaze direction from the start frame, and two presets handle front-facing and side-angle setups for an ai talking head shot.
- Custom Motion: You can direct gaze, gesture, posture, and energy in plain English when the audio alone does not carry the delivery you want.
- Long-Form Stability: A 30-minute talking video renders in one pass with the face holding steady, which suits a full course builder workflow.
- Cinematic Scenes: A verified Digital Twin can appear inside Avatar Shots, which places your face in Seedance 2.0 footage for clips up to 15 seconds.
- Developer Access: The API exposes Avatar V at $0.05 per second, so product teams can generate presenter footage inside their own apps.
What Could Be Better
- Avatar V does not run on photo-based looks, so you need at least one video look before it becomes available.
- Prompts have the lowest priority on this model, so flat or monotone audio produces a static avatar no matter what you type.
Verified Customer Results
Educator Anton Voroniuk saves 15.5 hours per week producing courses with his avatar and has reached more than 1 million students. On the enterprise side, Advantive cut content creation time by 50% and moved voice-over production from days to 2 to 3 hours for a team supporting 600+ employees.
Pricing
Avatar V renders cost 20 credits per minute on web plans, starting with Creator at $29 per month or $24 per month billed annually with 600 monthly credits. API usage runs $0.05 per second from a separate USD wallet.
This engine leads the comparison because it is the only option that learns your real motion and holds your exact likeness from a 30-second clip to a 30-minute course.
Digital Twin (on Avatar IV)
Before Avatar V shipped in April 2026, the Digital Twin ran on Avatar IV, and that engine remains a switchable option per video. I re-rendered the agent's twin on Avatar IV with the same script and look to isolate the difference.
Avatar IV maps audio into facial performance and captures body language, eyebrow raises, and pauses from your footage. Her hand gestures looked natural, and full-body framing worked well for a standing intro.
The difference showed up over time. By minute six of the training script, her smile read slightly wider than in her source footage. The Help Center lists the fix order for that issue: switch to Avatar V first, then Avatar IV, then re-record.
Key Strengths
- Prompt Responsiveness: Custom prompts like "talks excitedly" change gestures more on Avatar IV than on Avatar V, which helps for one-off creative cuts.
- Full-Body Motion: The twin supports full-body presence with clean or dynamic backgrounds, not only head-and-shoulders framing.
- Same Asset, Two Engines: You keep one twin and switch engines per video, so you can test both without re-recording.
- Speed or Quality Modes: Avatar IV offers a faster creation mode when turnaround matters more than polish.
Limitations
- Facial detail drifts more over long scripts than on Avatar V, which shows up most in teeth and smile width.
- Avatar IV renders cost the same 20 credits per minute as Avatar V, so you pay the same rate for less consistency.
- Fixing a drifting smile often means re-rendering, and every A-roll re-render consumes credits again.
Pricing
A Digital Twin slot comes free with one slot and up to 500 looks, and Avatar IV renders cost 20 credits per minute on paid plans.
The Avatar IV Digital Twin fits creative experiments that need prompt control, but it trails Avatar V on the likeness consistency that makes a twin worth building.
Photo Avatar
The agent's first instinct was the photo route because she hates filming. I uploaded her best headshot, and a talking version was ready before she finished picking a voice.
Photo Avatar brings a still image to life with generated expressions, head movement, and gestures. Avatar IV powers it by default, and it handles real people, illustrations, 3D renders, and non-human characters. That range makes it the most flexible option here for brand mascots and stylized presenters.
The trade-off is that the motion belongs to the model, not to you. Her real head tilt never appeared, and a wide smile in the source photo produced a mouth that looked too large on certain words. The Help Center recommends a subtle-smile photo for exactly that reason.
Key Strengths
- No Camera Needed: One photo starts the process, and uploading more photos gives the identity more reference material for future looks.
- Character Range: Avatar IV animates illustrations, 3D characters, and non-human designs that video-based avatars cannot represent.
- Prompt Control: Motion prompts on photo looks have more influence than on Avatar V, so you can ask for specific energy.
- Generous Slots: Free accounts get three photo avatar slots, and paid plans include unlimited slots with 500 looks each.
Limitations
- Avatar V does not support photo-based looks, so you cannot reach the most consistent engine without a video look.
- The model generates motion rather than learning yours, so close colleagues will notice your mannerisms are missing.
- Each photo avatar training costs 60 credits on paid plans, and each generated look costs 1 credit.
Pricing
Photo Avatar renders cost 20 credits per minute on Avatar IV or 3 credits per minute on Avatar III, with training at 60 credits per avatar on paid plans.
The photo route fits mascots, illustrated presenters, and quick tests, but it cannot match a video-trained twin for looking and moving like you.
Instant Avatar
The agent had an Instant Avatar from 2024, built from about two minutes of footage. Many longtime HeyGen users have one, and the name still appears in older tutorials and in some API error codes.
Instant Avatar was the original fast path to a custom video avatar on the platform. Today, the same "Clone a Real Person" flow creates a Digital Twin instead, from as little as 15 seconds of webcam footage plus a live consent clip.
Her old avatar still rendered and still lip-synced cleanly. However, the footage had a cut in the middle and inconsistent lighting, both of which current upload checks reject because they break the avatar's stability.
The bigger gap is engine access. Avatar V needs a clean, continuous video look, so an avatar trained on older footage often cannot reach the model that produced the best results in every test below.
Key Strengths
- Existing Assets Keep Working: Older Instant Avatars still render, so nothing breaks in a video library built on them.
- Low-Cost Testing: Rendering on Avatar III loops the footage and applies lip sync for 3 credits per minute, which works for script and timing tests.
- Familiar Workflow: The footage-to-avatar concept is identical to the Digital Twin, so upgrading takes one new recording.
Limitations
- Some older custom studio avatars do not support Avatar V, which locks them out of the most consistent engine.
- The Help Center's own guidance says redoing an older avatar with strong footage and the new Look Packs noticeably improves likeness.
- Looped footage on Avatar III repeats the same gesture cycle, which viewers spot in videos longer than a minute or two.
Pricing
Existing Instant Avatars render at 3 to 20 credits per minute depending on the engine, and new users create Digital Twins instead.
Instant Avatar earned its place in HeyGen history, but anyone still using one should spend 15 seconds recording a Digital Twin and render it on Avatar V.
Head-to-Head: 2-Minute Market Update
Same 2-minute script about local mortgage rates, same cloned voice, same office background look for all four avatars. I judged likeness, lip sync, and whether the agent's own clients could tell it was AI.
Avatar V Result
Avatar V reproduced her head tilt before each statistic and kept her eye line on the camera for the full two minutes. Lip shapes on "refinance" and "basis points" matched the cloned voice closely. Two of the three clients I showed it to assumed she had filmed it.
Digital Twin (on Avatar IV) Result
The Avatar IV twin captured her hand gestures well and looked convincing for the first minute. Around 1:20, her smile widened slightly on emphasized words. One client said it looked like her "on a different day."
Photo Avatar Result
The photo version delivered clean lip sync and pleasant generated motion. Without her real mannerisms, every client recognized her face but none recognized her delivery. It still worked as a social clip with captions from the subtitle generator.
Instant Avatar Result
Her 2024 Instant Avatar handled the script without errors on Avatar III. The looped source footage repeated a hand movement twice in two minutes, which one client noticed.
Winner: Avatar V. It was the only version clients mistook for a real recording.
Head-to-Head: 10-Minute Training Module
Same 10-minute onboarding script for new agents at her brokerage. This test measured identity drift over length, which is where avatar types separate most.
Avatar V Result
Teeth, jawline, and skin texture at minute nine matched minute one when I paused the frames side by side. The render held one take with no visible seams. For a training video her brokerage plans to reuse for a year, that stability matters more than any single frame.
Digital Twin (on Avatar IV) Result
Avatar IV finished the full 10 minutes in one pass and stayed watchable. Frame comparisons showed small shifts in smile width and cheek shape after minute six. Most viewers would not notice, but she did.
Photo Avatar Result
The photo avatar stayed stable in framing but grew repetitive in its generated gestures by minute four. For a long lesson, the lack of her real expressions made the content feel more like narration than teaching.
Instant Avatar Result
The legacy avatar rendered the module, but the looped footage made its gesture cycle obvious over 10 minutes. Re-recording as a Digital Twin was the clear next step.
Winner: Avatar V. It was the only avatar with no measurable drift across 10 minutes.
Head-to-Head: Spanish and Japanese Versions
Same 2-minute market update, translated into Spanish and Japanese with her cloned voice. I checked lip sync on unfamiliar phonemes and whether her voice still sounded like her.
Avatar V Result
The Japanese version handled fast syllable sequences without the mouth lagging behind the audio. Her cloned voice kept its warmth in both languages, and a bilingual client said the Spanish cut sounded like a natural second language rather than dubbing.
Digital Twin (on Avatar IV) Result
Spanish looked excellent on Avatar IV. In Japanese, a few rapid syllables showed mouth shapes that trailed the audio by a frame or two.
Photo Avatar Result
Lip sync stayed accurate in both languages, which is impressive from one photo. The generated motion did not change its rhythm for Japanese pacing, so delivery felt slightly mechanical. For quick localized social posts, it still beats ai dubbing over a static image.
Instant Avatar Result
The legacy avatar translated cleanly in Spanish. In Japanese, the looped source footage made the mismatch between gesture timing and speech rhythm more visible.
Winner: Avatar V. It kept her voice, face, and timing aligned in a language she does not speak.
Pricing Comparison
Avatar V and Avatar IV cost the same 20 credits per minute, so choosing Avatar IV saves nothing. Photo Avatar on Avatar III is the cheapest way to test scripts. Even so, the Creator plan's 600 credits buy about 30 minutes of Avatar V video with your real likeness.
Who Should Pick What
Pick Avatar V if you want viewers to see you, not an approximation of you. That covers real estate agents posting weekly market updates, founders sending investor messages, and creators publishing daily to TikTok without filming. It also covers L&D teams that need one presenter to stay consistent across a year of modules, with SCORM export on Business and SSO plus audit logs on Enterprise. Pair it with ai voice cloning for the closest match.
Pick a Digital Twin on Avatar IV if you already have a twin and want prompt-driven creative variations for a single campaign. Switching to avatar iv per video keeps one asset while you experiment.
Pick Photo Avatar if your presenter is a mascot, an illustration, or a 3D character with no video to record. It is also the fastest way to test an ai spokesperson concept before anyone sits in front of a camera.
Pick Instant Avatar if you have an existing library built on one and need to finish a project before re-recording. Plan to replace it with a Digital Twin soon after.
Final Verdict
A Digital Twin rendered on Avatar V is the right avatar for most people because it learns your real motion and keeps your exact face steady at any length. Avatar IV adds prompt control, Photo Avatar covers characters you cannot film, and Instant Avatar keeps older libraries running. Record 15 seconds on your webcam with HeyGen's free plan and compare it against your current avatar.
FAQs
Can I turn my Photo Avatar into a Digital Twin later?
Yes, but it means recording video, because a Digital Twin trains on footage of you speaking. You can keep your ai photo avatar and add a video look to the same identity. Once a video look exists, Avatar V becomes available for that avatar.
I have an old Instant Avatar. Do I lose my looks if I re-record?
No, your existing avatar and its videos stay in your account. A new Digital Twin creates a fresh avatar with its own 500-look capacity. HeyGen's Look Packs can then generate persona-based outfits and settings for the new twin in a few taps.
Does a Digital Twin need my voice too?
No, but a cloned voice gives the most natural match, and HeyGen's own guidance recommends one for twins. Public voices suit general business presenters, while designed voices fit stylized characters. Expressive source audio also makes Avatar V gestures livelier.
Can I use my avatar in a live sales call or support chat?
Yes, through LiveAvatar, which runs real-time conversational avatars for support, sales, and education. LiveAvatar is a separate platform with its own plans, so your HeyGen avatar slots do not transfer automatically. Your Digital Twin stays in your HeyGen account for pre-recorded video.
Which avatar type is cheapest to start with?
Photo Avatar costs nothing to start on the free plan, which includes three slots and one photo training per month in the mobile app. A Digital Twin is also free to create on the free plan with one slot. On paid plans, Avatar III renders at 3 credits per minute for low-cost script testing before a final Avatar V render.
Can I make a vertical video for Instagram or TikTok with any of these?
Yes, all four render in portrait format for short-form platforms. Avatar Shots adds cinematic Seedance 2.0 hooks for verified Digital Twins, and you can switch to a presenter shot for the main message. HeyGen's youtube shorts workflow handles the 9:16 framing and captions.
Can my avatar present the same course in five languages for my global team?
Yes, and your Digital Twin keeps your face and cloned voice in every version. Workday produces each training video in 10 to 15 languages and doubled capacity without adding headcount. The ai video translator also exports subtitles and an editable transcript for review.
Is my likeness safe when I create a Digital Twin?
Yes, Digital Twins require a live consent recording that matches the training footage, plus human moderation. The platform holds SOC 2 Type II, GDPR, and CCPA compliance, and enterprise data is never used to train models. A screen-recorded or AI-generated consent clip fails verification.
Greetings! My name is Ayesha Shaheryar. My words have helped millions over the past two years. As a HeyGen expert and a writer, I am here to introduce tips and tricks to edit your next video in no time.







