New in October 2026. Instant cloning now runs on the HeyGen Voice model and speaks through Text to Speech.
POST /v3/models/audio/voices with "mode": "instant" returns a voice_id that is usually ACTIVE within seconds. No training and no voice clone slot.
Set up with an agent
Not technical? Paste this into Claude Code, Codex, or Cursor and it walks you through every step on this page in your own terminal.Prompt for your agent
1. Create the voice
Upload the recording and pass itsasset_id, or send a public HTTPS URL or base64 audio. The call returns 202 Accepted.
Response
About the recording: only its first 3 minutes are used. A file can be up to 100 MiB, or 16 MB after decoding for inline base64. A URL must point at the audio file itself and download within 30 seconds; HeyGen fetches it after returning
202, and a fetch failure ends the voice FAILED with INVALID_AUDIO. A silent recording fails.
Each instant voice counts toward the workspace’s voice clone allowance. A request past it returns 400 resource_limit_reached; deleting a voice frees its place. An instant voice cannot be retrained. To change it, create a new one.
Idempotency-Key is optional. Reusing a key within 24 hours returns the first response, and a duplicate that arrives while the first is still being accepted returns 409 request_in_progress. Without the header, each retry creates a new voice.
2. Wait until it is ACTIVE
PollGET /v3/models/audio/voices/{voice_id} until status is ACTIVE or FAILED.
Response
A voice created without
language has no language field until HeyGen detects it. A failed voice stays in the list as FAILED until you delete it.
Polling is optional. Speech on a PENDING voice returns 409 voice_not_ready with a Retry-After header; wait that many seconds and send it again.
3. Generate speech
Pass the voice to Text to Speech.POST /v3/models/audio/tts returns one 44.1 kHz WAV; POST /v3/models/audio/tts/stream streams audio parts over Server-Sent Events from the same body.
model, voice_id, text, language, and expressiveness_boost (0.0–1.0, default 1.0), plus with_timestamps on the stream. seed, speed, pitch_shift, pitch_variance, and <break> tags work with professional voices only and return 400 invalid_parameter for an instant voice.
Languages
language takes one of these codes, both when you create a voice and when you generate speech: ar, be, bg, ca, cs, da, de, el, en, es, fa, fi, fr, he, hi, hr, hu, id, it, ja, ko, mk, ms, nl, pl, pt, ro, ru, sk, sl, sr, sv, ta, th, tl, tr, uk, vi, zh. Any other value returns 400 invalid_parameter.
Manage voices
GET /v3/models/audio/voices lists the workspace’s instant and professional voices, newest first; mode tells them apart. Set limit from 1 to 100 (default 10) and pass next_token as token while has_more is true.
DELETE /v3/models/audio/voices/{voice_id} deletes a voice. Deleting an instant voice that is still PENDING cancels it.
Pricing
Errors
Speech errors are on Text to Speech. The full catalog is in Error Codes.

