HeyGen Video

Today we're launching HeyGen Video, our first general-purpose video model, built for businesses that need production-quality video without production-level costs.

One model covers text to video (T2V), image to video (I2V) and reference to video (Ref2V): start from a prompt, a first frame, or your own reference images. Every clip comes back with sound (dialogue, ambience, and effects) generated in the same call.

In blind side-by-side tests, it performs alongside the top video models at a fraction of their price per second. That makes it practical for everyday business work: product demos, training, onboarding, and property walkthroughs.

Pricing starts at $0.01 per second through October (50% off the standard $0.02). It's live now on the HeyGen API and through OpenRouter, Runware, and ComfyUI.

Built on MiniMax H3 and post-trained by HeyGen.

Image to video

Cost

  • HeyGen$0.03/s
  • H3 Max Turbo$0.04/s
  • H3 Max$0.08/s
  • H3 Max balanced$0.08/s
  • Kling 3.0 Pro$0.168/s
  • Seedance 2.0$0.303/s
  • Veo 3.1$0.40/s

List price per second at 768p, with audio, before launch promotions.

Quality

  • HeyGen1000
  • Seedance 2.0955
  • H3 Max952
  • H3 Max Turbo920
  • H3 Max balanced860
  • Kling 3.0 Pro857
  • Veo 3.1744

Eval conducted on an internal eval set of Artificial Analysis Arena queries, with 4,800 votes. HeyGen's Elo is normalized to 1000 for simpler comparison.

Blind preference
True to image: 65% of 16 answersFollows prompt: 62% of 115 answersNo extras: 57% of 29 answersNatural motion: 63% of 75 answersTiming: 66% of 21 answersCamera motion: 47% of 21 answersPhysics: 59% of 72 answersConsistency: 43% of 34 answersDetail: 46% of 27 answersSound & voices: 57% of 59 answersMusic: 45% of 20 answers1True to image65%2Follows prompt62%3No extras57%4Natural motion63%5Timing66%6Camera motion47%7Physics59%8Consistency43%9Detail46%10Sound & voices57%11Music45%
  1. True to image65%
  2. Follows prompt62%
  3. No extras57%
  4. Natural motion63%
  5. Timing66%
  6. Camera motion47%
  7. Physics59%
  8. Consistency43%
  9. Detail46%
  10. Sound & voices57%
  11. Music45%

Eval conducted on an internal eval set of Artificial Analysis Arena queries, with 4,800 votes.

Speed · DiT inference time
  • HeyGen3.7 s
  • H3 Max Turbo4.1 s
  • H3 Max8.3 s

Inference time for one 10-second image-to-video clip. Captioner time is excluded, since it varies with the mode.

Quality vs price · 8 video models
6007008009001000$0.01$0.02$0.04$0.08$0.16$0.32$0.64Price per secondQuality (Arena Elo)H3 MaxH3 Max TurboH3 Max balancedSeedance 2.0Kling 3.0 ProVeo 3.1BorealHeyGen Video

Eval conducted on an internal eval set of Artificial Analysis Arena queries, with 4,800 votes.

The questions

Every word on screen is lettered by the model, inside the shot.

What if you could...Text to videoGenerated in 7 s
Generate video like this?Text to videoGenerated in 10 s

The race

One car at full speed, a prompt per shot.

OnboardText to videoGenerated in 8 s
Head-on passText to videoGenerated in 11 s
Driver close-upText to videoGenerated in 5 s
Side-on passText to videoGenerated in 7 s
TOP-TIER VIDEO QUALITY on the trackText to videoGenerated in 5 s

Realistic sound

A quiet platform, then the train. The sound is generated with the picture, in the same call.

The train arrivesText to videoGenerated in 17 s

Ten businesses

One idea per vertical, a second each in the film.

CaféText to videoGenerated in 17 s
Pharma labText to videoGenerated in 23 s
Real estateText to videoGenerated in 22 s
HospitalityText to videoGenerated in 22 s
Fashion retailText to videoGenerated in 14 s
FitnessText to videoGenerated in 8 s
AutomotiveText to videoGenerated in 16 s
TravelText to videoGenerated in 10 s
RestaurantText to videoGenerated in 10 s

Dawn to night

Street, farm, industry, art and the artists, each shot handing its motion to the next.

Dawn streetText to videoGenerated in 9 s
Farm at sunriseText to videoGenerated in 12 s
Textile millText to videoGenerated in 17 s
SteelworksText to videoGenerated in 18 s
GlassblowingText to videoGenerated in 10 s
PotteryText to videoGenerated in 14 s
Painter's studioText to videoGenerated in 16 s
GalleryText to videoGenerated in 15 s
Dance rehearsalText to videoGenerated in 11 s
CellistText to videoGenerated in 30 s
Fashion shootText to videoGenerated in 38 s
Construction at duskText to videoGenerated in 26 s
Harbour at duskText to videoGenerated in 25 s
Night marketText to videoGenerated in 8 s
Night tramText to videoGenerated in 12 s
TheatreText to videoGenerated in 8 s
City at nightText to videoGenerated in 14 s
4 a.m. bakeryText to videoGenerated in 11 s
LighthouseText to videoGenerated in 10 s
Port at nightText to videoGenerated in 14 s
Night highwayText to videoGenerated in 48 s

The close

The words are printed into the paper by the model.

HeyGen VideoText to videoGenerated in 15 s
Available via APIText to videoGenerated in 8 s