Skip to main content
HeyGen Video 1 generates the whole frame. Give it a prompt and it returns a clip of 5 to 15 seconds with its own audio track: dialogue, ambience and sound effects, synthesized in the same call. No separate speech pass, no lip-sync step, no avatar to pick. It is a different tool from the avatar rendering engines. Those animate a look you already own from a script you already wrote. This one invents the subject, the setting, the light and the sound at once, from a description.

What it is good at

Sound on for everything below. Every clip on this page is raw model output, with the audio it generated.

Scenes that hold together

Objects stay the same size and shape, stay where they were put, and obey physics. Things do not duplicate, vanish or reappear mid-shot.

Doing what you asked

It follows a brief literally rather than improvising. If you did not ask for dialogue, you do not get dialogue.

Picture and sound together

Room tone, effects and speech are generated with the image from the same prompt, so they match the space you described.

Holding an object you supply

Give it a photograph in reference_to_video and the colour, form and markings of that object survive into the shot.
It is at its best on short, contained shots: one subject, one place, one action, a locked or barely moving camera. It is least reliable on legible on-screen text, soft organic motion such as petals, paper and hair, and close hand work like assembling or operating equipment.

How to prompt it

Long prompts beat short ones. A few hundred to a few thousand characters gives materially better output, and there is no penalty for detail. The examples below each show the prompt that produced them.

The rules worth knowing

A body, a lens and an aperture set depth of field, grain and motion together. “Cinematic” and “high quality” do almost nothing. Shot on an ARRI Alexa Mini, 85mm at T2.0, one north-facing window as the only light changes the image.
Describing pores and fine lines gets you part way. Adding no retouching, no beauty filter, no glamour lighting gets you the rest. Without the second clause, faces drift toward a retouched look. The same applies to colour: ask for muted colour, then rule out the default with no teal-orange grade.
Plain clothing and equipment will otherwise pick up invented branding. Add no logos, brand names, printed words or badges anywhere in frame.
One short capitalized word renders reliably. Longer strings, multi-word titles and labelled diagrams come back as plausible-looking gibberish. Suppress lettering in the prompt and composite real type afterwards.
A static camera with the subject wearing or holding an object is the most reliable composition available. Objects worn on the body render best of all. Close hand work degrades fastest: if a procedure needs explaining, let the speech carry the steps and keep the hands still.
When an action involves one thing meeting another, name the contact point and say it holds. Without that, the gesture often lands near the target rather than on it.
Audio comes from the same prompt, so direct it. Name the room tone, the specific effects and their distance. Say no music when you do not want a bed, because one may otherwise appear. Written dialogue fits at roughly 2.5 words per second.
Set seed yourself and the same prompt returns the same clip, so you can change one clause at a time. Omit it and the server picks a random one, so two identical requests give you different videos.

Twenty styles

Style is a prompt variable. The reliable way to get a real one is to name an actual production process and the physical marks it leaves, restrict the palette, and rule out the smooth digital default.

Brand moments

A five-second identity piece, rendered in a medium chosen for the sector rather than a generic logo spin. Short single-word marks render reliably, which is what makes these work. A full lockup with a tagline still needs type added afterwards.
The brands below are invented for this page. Any resemblance to a real company is coincidental.

Build with it

Generate a video

Submission returns 202:
Send an Idempotency-Key header to retry safely: a retry with the same key within 24 hours returns the original response, and a concurrent duplicate returns 409.

Retrieve the result

Poll GET /v3/models/videos/{video_id}. pending means queued and processing means running. The terminal states are completed, failed and cancelled.
video_url is a signed link: poll again to refresh it. Failed and cancelled jobs carry failure_code (generation_failed or generation_cancelled) and failure_message. An unknown ID, or one from another workspace, returns 404. The same job is also readable through GET /v3/videos/{video_id}, alongside your other videos, with video_page_url and a title taken from the first 64 characters of the prompt.

Callbacks

Pass callback_url on the create request to be notified when a job finishes, and callback_id to correlate the delivery. Send callback_id alone to deliver to the webhook endpoints registered for your workspace. One event fires per job, at the terminal status, and a callback is attempted once. Treat it as a latency optimisation and keep polling as the fallback for anything you cannot afford to miss.

Modes

mode selects how the model is conditioned. Each mode has its own fields; a field from another mode returns 400 when it carries a value (an empty list or null passes). The short forms t2v, i2v and ref2va are accepted as aliases. New integrations should use the long names, which is what validation errors return. image_to_video treats the supplied image as the literal first frame, so the clip opens exactly as that still looks. To place a product or piece of equipment into a scene of your own, use reference_to_video and describe the scene around it.

References and prompt labels

In reference_to_video, list order becomes the label you address in the prompt. The first entry of reference_images is <Picture 1>, the second is <Picture 2>, the first entry of reference_videos is <Video 1>, and so on. Images, videos and audio are numbered independently. The API does not rewrite these labels. Prompt enhancement can reword the rest of the prompt; set it to disabled to keep your text as written. Each reference accepts an HTTPS URL, an uploaded asset ID, or inline base64.
URLs are fetched server side with SSRF checks and staged before dispatch. The fetcher does not follow redirects, so upload anything you do not control through POST /v3/assets and pass the returned asset_id.
Uploaded assets must belong to the calling workspace and are reusable across requests.

Prompt enhancement

Before generation, the prompt passes through an enhancement step that expands it for the model. prompt_enhancement picks how:
Use disabled when you have already written a long, fully specified prompt and want it followed word for word.

Parameters

The schema is strict: unknown fields are rejected rather than ignored.

Output

resolution names a size class and aspect_ratio sets the shape. When you set aspect_ratio explicitly, the two resolve to a fixed pixel size:

Default aspect ratio

text_to_video has nothing to take a shape from and defaults to 16:9. reference_to_video defaults to adaptive: the aspect ratio of the first reference image, or the first reference video when there are no images. Pass one of the six ratios to override it. image_to_video always follows the first frame, including its EXIF orientation. To change the output shape, crop the image before uploading it. In both adaptive cases the output follows the source ratio, scaled to the short edge of the resolution with each side rounded to a multiple of 32. A 1280 × 852 reference at 480p returns 736 × 480. The completed job reports the actual width and height, and aspect_ratio as the reduced pixel ratio, for example 23:15.

Seeds

Set seed yourself to make a shot repeatable: the same prompt and the same seed, submitted in succession, return an identical file. Hold the seed and change one clause at a time to iterate. When seed is omitted the server picks a random one, so two identical requests return different videos. The completed job reports the seed it used as seed. Treat a seed as reproducible within a deployment rather than as a permanent handle on one render.

Errors

Errors return {"error": {"code", "message", "param", "doc_url"}}. Validation errors use code invalid_parameter and set param to the field. This route runs on paid API keys.

Specs at a glance

Get an API key

Create a key and check your usage.

HeyGen Avatar

Animate a look you own from a script, with Avatar V, IV and III.