background leftbackground right

What’s new at HeyGen: September 2026

Holly Xiao
Written byHolly Xiao
Last UpdatedOctober 1st, 2026
What's new at HeyGen September 2026, with the HeyGen logo and a plus icon.
Create AI videos, starring you in 177+ languages and dialects.
Get started for free
Summary

September was built for developers. HeyGen shipped an open-source real-time avatar stack powered by GPT-Live-1, launched the Code2Video Benchmark for evaluating AI-generated motion graphics, and brought Professional Voice Clone to the API.

September was for builders. None of this month's releases is a new button in the HeyGen editor. Each one is a piece you use to build your own product: a working stack to start from, a benchmark to measure against, and a voice model you can call from code. Here is what shipped, what it does, and where to get it.

An open-source real-time avatar stack, built with OpenAI’s GPT-Live-1

A chatbot that answers correctly in a text block doesn't pass anymore. The avatar hears you, responds, and shows its work on screen while it does it.

On September 10 we open-sourced a reference integration that does exactly that, built with OpenAI. liveavatar-gpt-live-demos wires three systems together: GPT-Live-1, OpenAI's full-duplex speech-to-speech model, drives a HeyGen LiveAvatar, and the model's tool calls land on screen as animated HyperFrames overlays. You talk, the avatar talks back, and what it does appears next to it.

Building this from nothing means getting speech in and out, a real-time model, avatar rendering, tool calling, and interface state all working together. The demo is that wiring, written to be read.

  • Start from three working demos, not an empty repo. The repo ships three demos: a language tutor, a poker coach, and a customer support agent. The tutor, for example, teaches a word aloud, puts a term card on screen as it says it, and every few words shrinks to the corner to review what you've learned so far. Three commands (pnpm install, pnpm run setup, pnpm dev) and you're talking to it.
  • Swap the persona, keep the face. The persona is two markdown files; adding a new on-screen visual is three small edits. Change those and the tutor becomes a support agent that looks up an order and shows the result, a sales assistant that talks through a configuration while displaying it, or a trainer that walks someone through a workflow one step at a time.
  • Generate the visuals live. A second repo, liveavatar-hyperframes-demo, has the avatar render motion graphics on request. Say "animate three orange cats" and a background Claude Code session writes a HyperFrames composition that fades in over the avatar's video. Generation takes 30 seconds to two minutes, so the avatar acknowledges the request and keeps talking while it works.

Both repos are MIT licensed. You'll need a LiveAvatar API key, plus an OpenAI key with GPT-Live access for the first demo. They are starters, not deployments: add auth before you expose one publicly, and read the hardening list in the repo's architecture doc.

Code2Video: A benchmark for video, not just code

Code that compiles into a video doesn't mean it’s a good video. The video has to look intentional, paced, and coherent, and the benchmark exists to measure exactly that.

On September 18, HeyGen Research released the Code2Video Benchmark with Kaggle. It measures how well modelsturns a creative brief into motion-graphics code that HyperFrames renders into a finished video. Models are good at writing code that runs. They are much less good at timing, composition, and craft, and until now there was no shared way to measure the gap.

  • 168 briefs, each a beat of a real launch video. Hooks, problem statements, product intros, key features, benefits, social proof, calls to action, and brand outros, each with a human-made reference to measure against.
  • Five axes make up the score. Engagement, prompt-intent, composition, temporal, and craft are judged separately, so you can see where a model is strong and where it falls apart.
  • A judge trained on human preference. General-purpose vision-language models agree with human raters about 75% of the time on this task, and they carry position and self-preference biases. Our purpose-built judge model reaches 82%, returns a verdict in under a second, and predicts, for any two videos, which one people would prefer.
  • An open, audited leaderboard. Every model run executed on Kaggle, which keeps the raw traces auditable and will add new models as they launch. Rankings are Elo, so you can watch a model move from #14 to #10 even while every model still loses to the human reference.

Early findings: the top four models sit within 12 Elo of each other, open-weight models land within 50 Elo of first place, and following the brief is nearly solved while craft is not. Motion is the weak axis.

If you are building agents that produce video, this is your target. The full report shows the common failure modes side by side, and running your own models against the same judge on Kaggle is coming next.

Professional Voice Clone, now in the API

An avatar that looks like someone but sounds like a stock voice doesn't pass. In September, Professional Voice Clone became available through the HeyGen API: a dedicated voice adapter trained on the HeyGen Voice model from at least 20 minutes of one speaker's recordings. It is the highest-fidelity clone we offer, and it now fits inside a pipeline.

  • Train once, then generate on demand. Upload 1 to 10 recordings with the Assets API, create the voice with POST /v3/models/audio/voices, poll until it is ACTIVE, then call POST /v3/models/audio/tts for a finished WAV or /v3/models/audio/tts/stream for audio as it is generated.
  • Keep the same voice across releases. A creator can localize or update narration without another recording session. A training team can keep every course in the same instructor's voice.
  • Retrain without starting over. Send new recordings with the existing voice_id and the ID, name, language, and slot stay the same. If retraining fails, the previous voice stays active.
  • Not a developer? Let your coding agent do it. The docs include a prompt you can paste into Claude Code, Codex, or Cursor that walks you through training a clone of your own voice and hearing it speak, one step at a time.

Each voice occupies a slot purchased on the API usage page, each slot includes five trainings per billing period with the first training counted, and synthesis bills 0.6 API credits per generated minute. Voices belong to the workspace of the API key that created them, and deleting a voice frees its slot.

Start building

Everything above is live today. The demos are on GitHub, the benchmark is on Kaggle, and Professional Voice Clone is in the API docs. If you are new to the HeyGen API, start at developers.heygen.com and show us what you build.

About

Meet Holly Xiao, Head of Marketing at HeyGen. With deep expertise in product and growth marketing, Holly has led marketing teams at Drift, Envoy, and Canvas, crafting narratives that fuel business growth through clear positioning and storytelling. At HeyGen, she’s helping redefine how businesses use AI-powered video to scale enterprise communication and engagement.


Continue Reading

Latest blog posts related to What’s new at HeyGen: September 2026.

Browse All

Start creating videos with AI

See how businesses like yours scale content creation and drive growth with the most innovative AI video.

CTA background