Make Photo Sing features
Pixel-accurate AI lip sync engine
Every syllable, breath, and beat in your favorite song lines up with the face. The AI lip sync engine synchronizes mouth shapes at the phoneme level, so each performance looks real, not robotic. Even held notes and runs land clean.

Smart face detection on any portrait
Drop in a selfie, a pet picture, a cartoon character, or a vintage portrait. The image-to-video pipeline reads facial features, locks framing, and reconstructs subtle head motion so each animation holds together across long clips and close-up shots.

Bring your song or AI-generated voice
Upload an MP3 or WAV file containing your full-length track. Prefer a bespoke narration? Create one using the AI voice generator and send it straight to the singing photo. Both options use the same advanced engine to ensure consistent quality.

Expressive facial animation, without a robotic feel
Pick a mood, and the AI face follows. Subtle blinks, smiles, brow lifts, and head tilts arrive automatically, so a ballad reads tender and a hype track reads fired up. No frame-by-frame work or animation timeline to fight with later.

Built-in captions and vertical exports
Burn in word-by-word captions with the subtitle generator and export in 9:16 for Reels, 1:1 for feeds or 16:9 for YouTube. Safe text placement keeps your hook within the frame on every platform, ensuring it is visible at first glance whenever you post.


Static selfies do not stop people scrolling. Turn any photo into a singing image based on the latest trend, post a few variations in succession, and let the format and timing work for you across social feeds.

A standard e-card can feel like a chore. Make Grandma sing ‘Happy Birthday’ in her own voice, send your best friend a version featuring your shared joke, or use a bespoke song for a wedding toast that is both fun and engaging.

Pets are pure social fuel. Upload a clear photo of your dog, cat, or hamster, drop in a hilarious song to sing your favorite tunes, and watch I Love Happy Cats-style cuteness across every feed you post to.

Animate faces from illustrations, channel mascots, AI-generated cartoon characters or your virtual identity without rigging or motion capture. The animator gives every avatar a convincing on-stage performance, post after post, without taking up filming time in your diary.

Animate faces from illustrations, channel mascots, AI-generated cartoon characters, or your virtual identity without rigging or motion capture. The animator gives every avatar a real on-stage performance, post after post, with no filming time on your calendar.

Have historical figures sing their dates, animate book characters as they read lyrics, or use the AI video translator to translate a song into more than 177 languages, so the same lesson is ready to share across different languages and classrooms.
How the Make Photo Sing tool works
Make your photo sing in four steps. Animate photos without filming, editing experience or plug-ins.
Choose a clear, front-facing image of a person, pet or character. The AI identifies facial features.
Add an MP3 or WAV file, choose a song from your library, or paste in lyrics for the AI to sing.
Choose an emotion to shape the performance and select the portrait or square format you plan to post.
Render the AI-generated photo, preview the singing result, and download an MP4 ready for any platform.
It means turning a still image into a short video in which the face performs along to an audio track. The AI automatically detects facial features, maps phonemes to mouth shapes, then adds blinks and head movements to bring the photo to life as it sings your chosen song.
Sign up online for a free HeyGen account, upload a clear, front-facing portrait, add an audio file and select ‘Generate’. The Free plan includes short clips with a watermark, whilst a paid plan provides longer renders and HD output.
Clear, well-lit, front-facing headshot images work best. Avoid significant obstructions, sunglasses, extreme side angles and low-resolution images. Photographs of pets, cartoons and illustrations all work, provided the face is visible and centred in the frame.
Yes. Add any MP3 or WAV file, including original tracks, covers, voice memos or instrumentals. The AI analyses the vocals, beat and tone, then maps the performance to match the rhythm. Confirm that you hold the rights to any commercial music before posting.
Yes. HeyGen handle multiple languages with natural pronunciation that matches the rhythm of each language. Run the voice through the AI video translator into Spanish, Hindi, Mandarin or any other language, and the lip-synching follows.
Yes. Open the clip in the AI video editor to trim it, adjust expression intensity, replace the audio, generate alternative takes, or add captions, background music and brand colours before exporting the final cut.
Yes. If the image has a recognisable face, the AI animates it. Dogs, cats, illustrated cartoon characters, anime portraits and AI-generated avatars all work. Singing pets and cartoons are amongst the most popular formats on the platform.
The tool animates one face at a time for the cleanest result. Crop the image to show the person you want as the singer, or generate separate clips for each face and stitch them together in post-production to create a group scene.
Yes, you own the outputs. Please confirm that you hold the rights to any third-party music, photographs or voices you upload. HeyGen serve more than 85,000 businesses, ranging from solo creators such as Anton Voroniuk who reach over 1 million students, to global brands.
Most short clips render in a few minutes. The AI-powered output is an MP4 file, available in vertical 9:16, square 1:1 for feeds and widescreen 16:9 for YouTube. Captions can be embedded for autoplay or kept separate for reuse across channels.
Yes. If you hold the necessary rights to the photo, voice and music, you can use the finished clip in commercial campaigns. For campaign formats and creative variations created from product information, images or scripts, use AI video adverts.
Explore more AI-powered tools
Bring any photo to life with hyper-realistic voice and movement using Avatar IV.
Make any photo sing with AI in minutes. Upload a picture, add a song, then share it on TikTok and Reels.
