Skip to main content
New in October 2026. Instant cloning now runs on the HeyGen Voice model and speaks through Text to Speech.
One recording in, a voice out. POST /v3/models/audio/voices with "mode": "instant" returns a voice_id that is usually ACTIVE within seconds. No training and no voice clone slot.
  1. Create the voice
  2. Wait until it is ACTIVE
  3. Generate speech

Set up with an agent

Not technical? Paste this into Claude Code, Codex, or Cursor and it walks you through every step on this page in your own terminal.
Prompt for your agent

1. Create the voice

Upload the recording and pass its asset_id, or send a public HTTPS URL or base64 audio. The call returns 202 Accepted.
Response
About the recording: only its first 3 minutes are used. A file can be up to 100 MiB, or 16 MB after decoding for inline base64. A URL must point at the audio file itself and download within 30 seconds; HeyGen fetches it after returning 202, and a fetch failure ends the voice FAILED with INVALID_AUDIO. A silent recording fails. Each instant voice counts toward the workspace’s voice clone allowance. A request past it returns 400 resource_limit_reached; deleting a voice frees its place. An instant voice cannot be retrained. To change it, create a new one. Idempotency-Key is optional. Reusing a key within 24 hours returns the first response, and a duplicate that arrives while the first is still being accepted returns 409 request_in_progress. Without the header, each retry creates a new voice.

2. Wait until it is ACTIVE

Poll GET /v3/models/audio/voices/{voice_id} until status is ACTIVE or FAILED.
Response
A voice created without language has no language field until HeyGen detects it. A failed voice stays in the list as FAILED until you delete it. Polling is optional. Speech on a PENDING voice returns 409 voice_not_ready with a Retry-After header; wait that many seconds and send it again.

3. Generate speech

Pass the voice to Text to Speech. POST /v3/models/audio/tts returns one 44.1 kHz WAV; POST /v3/models/audio/tts/stream streams audio parts over Server-Sent Events from the same body.
An instant voice takes model, voice_id, text, language, and expressiveness_boost (0.0–1.0, default 1.0), plus with_timestamps on the stream. seed, speed, pitch_shift, pitch_variance, and <break> tags work with professional voices only and return 400 invalid_parameter for an instant voice.

Languages

language takes one of these codes, both when you create a voice and when you generate speech: ar, be, bg, ca, cs, da, de, el, en, es, fa, fi, fr, he, hi, hr, hu, id, it, ja, ko, mk, ms, nl, pl, pt, ro, ru, sk, sl, sr, sv, ta, th, tl, tr, uk, vi, zh. Any other value returns 400 invalid_parameter.

Manage voices

GET /v3/models/audio/voices lists the workspace’s instant and professional voices, newest first; mode tells them apart. Set limit from 1 to 100 (default 10) and pass next_token as token while has_more is true. DELETE /v3/models/audio/voices/{voice_id} deletes a voice. Deleting an instant voice that is still PENDING cancels it.

Pricing

Errors

Speech errors are on Text to Speech. The full catalog is in Error Codes.