background leftbackground right

How to Write Prompts for AI Video: Guide and Examples

Ayesha Shaheryar
Written byAyesha Shaheryar
Last UpdatedOctober 9th, 2026
How to Write Prompts for AI Video: Guide and Examples
Create AI videos, starring you in 177+ languages and dialects.
Get started for free
Summary

How to write prompts for AI video: the 5-part structure, style paragraphs, scene scripts, and copy-ready examples for agents, scenes, and avatars. Try it free.

TL;DR: The fastest path is writing a creative brief for HeyGen's Video Agent, not a list of adjectives, because the agent paces scenes to the duration you state, follows a pasted script scene by scene, and applies a style paragraph literally to every graphic. One structured prompt gets about 90% of the way to a finished cut.

  • Templates are the better call when you need the same layout every week and don't want to write prompts at all.
  • Cinematic models like Veo 3.1 and Kling 3.0 suit single atmospheric shots where camera language matters more than a script.
  • Drafting your prompt in ChatGPT or Gemini first helps when you're starting from a rough idea rather than a finished outline.

Most AI video prompt guides tell you to add "cinematic, 4K, ultra-realistic, masterpiece." HeyGen's video model documentation says generic words like "cinematic" do almost nothing, and that naming a specific camera and lens changes the result far more.

That gap explains why so many first attempts look like stock footage. The model fills every decision you leave open, so vague prompts produce average videos.

This guide covers how to write prompts for AI video in five parts, the three different prompt types most guides blur together, and copy-ready examples. The developer documentation describes the prompt as the whole interface to its Video Agent, and the steps below follow that structure.

How to Write AI Video Prompts with HeyGen

HeyGen homepage screenshot

1. Lead with format, length, and orientation

Open with one sentence that sets the container: "Make a 45-second portrait video" or "Make a 3-minute horizontal explainer." The agent paces every scene to that duration.

Orientation matters for layout, since portrait suits mobile and social while horizontal suits presentations. This first line takes 30 seconds to write and fixes the most common pacing problem.

2. State the audience, the goal, and the one message

Add who's watching and what they should do after: "for new sales hires, so they can explain our pricing tiers on their first call." Then name the single message the video has to land.

A prompt with one message produces a focused video. A prompt with five messages produces five rushed scenes. Spend 2 minutes here, because it decides what the agent cuts.

3. Paste the script and ask for scene-by-scene control

If you have a script, paste the whole thing and add "Follow the script below exactly, scene by scene." Label scenes with timestamps and say what appears on screen: presenter, motion graphic, or B-roll.

Without a script, the agent writes one for you, which works for drafts. For text to video projects where wording matters, such as legal or product claims, always paste your own.

4. Add a style paragraph with exact colors

Write five or six sentences that give the look a name, the exact palette in hex codes, the art direction, how elements move, the transitions, and one closing line for the vibe. The agent builds graphics in code with HyperFrames, so it follows this paragraph literally.

Skip words like "modern" and "clean" without detail. "Ivory #F7F4EE background, charcoal serif headlines, numbers count up gently" gives the agent rules instead of moods. This takes about 5 minutes the first time and seconds after you save it.

5. Attach references, review the plan, and revise one scene at a time

Attach up to 20 files, such as slides, product images, PDFs, or audio, and say how to use each one. Switch to chat mode to see a scene-by-scene plan before anything renders.

After the first cut, fix single scenes ("make scene three's chart green") instead of rewriting the prompt. HeyGen's guidance sums it up as prompt for the 90, edit for the 10. The Video Agent prompt guide on the blog has more worked agent examples.

The Three Prompt Types Most Guides Blur Together

None of the ranking prompt guides separate the three jobs a prompt can do, and each one follows different rules. A prompt that works for one fails for the others.

Loading embed content...

The scene row comes from HeyGen's video model documentation, which lists long on-screen text, soft organic motion like hair and petals, and close hand work as its weakest areas. Spoken dialogue fits at about 2.5 words per second of clip, so a 10-second shot holds about 25 words.

The avatar row reflects how Avatar V weighs inputs: audio first, the source image's expression second, and the text prompt third. Flat audio gives a flat performance no matter what you type, which is why voice choice matters more than motion prompts.

Prompt Examples You Can Copy

Agent brief for a product explainer

Make a 60-second horizontal explainer for small business owners about our new invoicing app. Goal: they understand it sends reminders automatically.

One presenter on camera for the intro and outro, motion graphics for the three features, and a call to action to start a free trial.Style: Ledger Clean. White #FFFFFF, Ink #1B1F24, Mint #2BB673, Amber #F5A623. Flat icons with thin outlines and generous whitespace.

Numbers count up, icons slide in from the left. Transitions are quick horizontal wipes. Calm, competent, and friendly.

Scene prompt for B-roll

Medium shot of a ceramic coffee mug on a walnut desk beside an open laptop, morning light from a window on the left. Slow push-in. Shot on an ARRI Alexa Mini with a 50mm lens, shallow depth of field. Sound: quiet room tone and a distant keyboard. No music.

Script-led agent prompt for a training module

Make a 2-minute horizontal training video for warehouse staff. Follow the script below exactly, scene by scene. Use the presenter on camera for scenes 1 and 5 and motion graphics for scenes 2 to 4.Scene 1 (0-15s), presenter: "Every lift starts before you bend."Scene 2 (15-45s), graphic of feet shoulder-width apart: "Plant your feet shoulder-width apart..."

Motion note for an avatar video

Warm and confident. Look into the camera on the opening line, glance down at the notes on "here's the number," then return to camera. Keep hand movement below the chest.

Quick Prompt vs. Full Brief

Depending on how much the output matters, use one of two prompt lengths.

The 30-second prompt

One sentence with format, length, audience, and topic: "Make a 30-second portrait video for LinkedIn explaining why our team switched to a four-day week." The agent decides the script, presenter, and look. Use it for drafts, idea testing, and internal clips.

The full brief

Format line, audience and goal, pasted script with scene labels, style paragraph, and attached references. Use it for anything customer-facing or anything you'll re-use as a template. Save the style paragraph once and paste it into every future brief for a consistent channel look.

Common Mistakes to Avoid

Stacking adjectives instead of decisions

"Epic, cinematic, stunning, high-quality" gives the model nothing to act on. Replace each adjective with a decision: a lens, a hex code, a transition, a camera move.

Forgetting to say "no music"

When a model generates sound with picture, it may add music you didn't ask for. Say "no music" and name the room tone and effects you want, with distances, such as "a door closing in the next room."

Asking for long on-screen text

Generated paragraphs of on-screen text often come out garbled. Keep generated text to a word or two, and add longer captions, lists, and lower thirds in the editor after rendering.

Rewriting the whole prompt to fix one scene

A full rewrite reshuffles everything that already worked. Use chat-mode revisions on the single scene and keep the rest locked.

Other Ways to Direct AI Video

Templates

Template-based editors, including HeyGen's own template library, give you a fixed layout to fill with your script.

  • Pros: no prompting skill needed; consistent layouts every time; fast for recurring formats; easy to hand off to teammates.
  • Cons: you're limited to the layouts that exist, so a unique visual idea needs a prompt anyway; templates can look familiar to viewers who've seen the same design elsewhere.

Cinematic models like Veo 3.1 and Kling 3.0

Google Veo 3.1 / Flow homepage screenshotKling 3.0 homepage screenshot

Google's Veo 3.1 and Kuaishou's Kling 3.0 are video generation models that turn a scene prompt into a short shot with native audio.

  • Pros: strong camera language and lighting; native sound in the same generation; good for atmospheric B-roll; flexible for creative experiments.
  • Cons: clips top out at 8 seconds on Veo 3.1 and 15 seconds on Kling 3.0, so a full explainer means stitching many shots; no script-faithful presenter for teaching or sales content.

Drafting prompts in ChatGPT or Gemini

Many users draft their prompt in a chatbot, then paste it into the video tool.

  • Pros: turns a rough idea into a structured brief; helps with script length and tone; free to start; good for brainstorming angles.
  • Cons: chatbots tend to write adjective-heavy style notes, so you still need to replace moods with specifics; they don't know your brand's hex codes or terms unless you provide them.

Frequently Asked Questions

How long should an AI video prompt be?

As long as the decisions you need to make. A 30-second social clip can work from one sentence, while a script-led explainer with a style paragraph often runs 300 to 600 words. Video Agent accepts prompts up to 10,000 characters through the API, so length is rarely the limit.

Should I write the script or let the AI write it?

Let the AI write it for drafts and idea testing. Write it yourself, or edit the AI's draft, for anything with claims, prices, or regulated language. Pasting your script with "follow exactly" keeps every word under your control, which matters when legal or compliance teams review the final cut.

Do I need to describe the avatar in the prompt?

Only if you want the agent to choose. To keep one presenter across a series, pin a specific avatar instead of describing one, because a description can match a different face each time. Your own Digital Twin works the same way with AI voice cloning for a consistent voice.

Why does my video look generic even with a detailed prompt?

Usually because the detail is in adjectives rather than decisions. Check the style paragraph for hex codes, motion rules, and transitions, and check scene prompts for a named lens and camera move. Attaching your own product images also replaces stock-looking visuals.

Can I reuse a prompt for a weekly series?

Yes. Save the format line and style paragraph as a template, then swap only the audience line and script each week. Pinning the avatar and voice keeps the series consistent, and the agent's output stays close from episode to episode.

Conclusion

To write prompts for AI video that land on the first cut, state the length, name the audience and message, paste the script, add a style paragraph with hex codes, and revise single scenes. Test it on the Free plan, then move to Creator at $29 per month.

About

Greetings! My name is Ayesha Shaheryar. My words have helped millions over the past two years. As a HeyGen expert and a writer, I am here to introduce tips and tricks to edit your next video in no time.


Continue Reading

Latest blog posts related to How to Write Prompts for AI Video: Guide and Examples.

Browse All

Start creating videos with AI

See how businesses like yours scale content creation and drive growth with the most innovative AI video.

CTA background