> ## Documentation Index
> Fetch the complete documentation index at: https://heygen-1fa696a7.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Instant Voice Clone

> Clone a voice from one recording on the HeyGen Voice model with POST /v3/models/audio/voices, then generate completed or streaming speech.

<img className="w-full h-44 object-cover rounded-xl" src="https://mintcdn.com/heygen-1fa696a7/hfMXXwJzjE7vBSYZ/images/theme/research-2.webp?fit=max&auto=format&n=hfMXXwJzjE7vBSYZ&q=85&s=d6ee10585b27ddaf8137d6ed67791708" alt="" noZoom width="1400" height="788" data-path="images/theme/research-2.webp" />

<Note>
  **New in October 2026.** Instant cloning now runs on the [HeyGen Voice](/docs/models/heygen-voice) model and speaks through [Text to Speech](/docs/voices/heygen-voice-speech).
</Note>

One recording in, a voice out. [`POST /v3/models/audio/voices`](/reference/create-or-retrain-an-audio-voice) with `"mode": "instant"` returns a `voice_id` that is usually `ACTIVE` within seconds. No training and no voice clone slot.

1. [Create the voice](#1-create-the-voice)
2. [Wait until it is `ACTIVE`](#2-wait-until-it-is-active)
3. [Generate speech](#3-generate-speech)

## Set up with an agent

Not technical? Paste this into [Claude Code](https://claude.com/claude-code), [Codex](https://openai.com/codex/), or Cursor and it walks you through every step on this page in your own terminal.

```text Prompt for your agent wrap theme={null}
I want to make a HeyGen instant voice clone from one recording of my voice and hear it speak. I am not a developer, so run the commands for me where you can, explain each step in one plain sentence, and when I have to do something myself tell me exactly what and wait for me.

Read these first and follow them exactly:
- https://developers.heygen.com/docs/voices/heygen-voice-instant-clone.md
- https://developers.heygen.com/docs/upload-assets.md
- https://developers.heygen.com/docs/voices/heygen-voice-speech.md
- https://developers.heygen.com/docs/for-ai-agents.md (rules for agents)

Then take me through this, checking each step before the next:
1. API key. Look for HEYGEN_API_KEY in my environment. If it is missing, send me to https://app.heygen.com/developers/api to create one and have me export it in my terminal myself. Never ask me to paste the key into this chat. Confirm it works with GET /v3/users/me.
2. Recording. Ask me for one audio file of my voice. Check it with ffprobe (install it if needed): only the first 3 minutes are used, so 1 to 3 minutes of clear speech is ideal. Warn me about long silence or background noise. Upload it with POST /v3/assets and keep the asset_id.
3. Clone. Ask me for a display name, then call POST /v3/models/audio/voices with mode instant, the asset_id and an Idempotency-Key. Poll GET /v3/models/audio/voices/{voice_id} until status is ACTIVE. If it is FAILED, explain failure_reason and what to change.
4. Speak. Ask me for a few sentences and their language, then call POST /v3/models/audio/tts with model heygen-voice-1, the voice_id, text and language. Download the WAV and open it so I can listen. Show me how to change the text and run it again. Never send seed, speed or pitch for an instant voice.

If a request fails, look the error code up in the Errors tables on those pages and tell me the fix. Never invent keys, IDs, or URLs; if something is missing, ask me for that one thing.
```

## 1. Create the voice

[Upload the recording](/docs/upload-assets) and pass its `asset_id`, or send a public HTTPS URL or base64 audio. The call returns `202 Accepted`.

```bash theme={null}
curl -X POST "https://api.heygen.com/v3/models/audio/voices" \
  -H "X-Api-Key: $HEYGEN_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: customer-instant-voice-2026-10-09" \
  -d '{
    "mode": "instant",
    "name": "Customer narrator",
    "language": "en",
    "audio": [
      { "type": "asset_id", "asset_id": "'"$ASSET_ID"'" }
    ]
  }'
```

```json Response theme={null}
{
  "data": {
    "voice_id": "7c2a9e4f1b3d4e6a8f0c2d4e6a8b1234"
  }
}
```

| Field | Type | Required | Description |
| - | - | - | - |
| `mode` | string | Yes | `instant`, lowercase. |
| `name` | string | Yes | Display name, 1–64 characters. |
| `audio` | array | Yes | Exactly one recording: `{ "type": "asset_id", "asset_id" }`, `{ "type": "url", "url" }`, or `{ "type": "base64", "media_type", "data" }`. |
| `language` | string | No | Language code of the recording, such as `en`. See [Languages](#languages). Detected from the recording when omitted. |
| `similarity` | string | No | How closely the voice follows the recording: `medium`, `high`, `extra_high`, or `max`. Default `extra_high`. |
| `remove_background_noise` | boolean | No | Clean the recording before cloning. Default `true`. |
| `gender` | string | No | `male` or `female`. Detected when omitted. |
| `age` | string | No | `young` or `old`. Detected when omitted. |

About the recording: only its first 3 minutes are used. A file can be up to 100 MiB, or 16 MB after decoding for inline base64. A URL must point at the audio file itself and download within 30 seconds; HeyGen fetches it after returning `202`, and a fetch failure ends the voice `FAILED` with `INVALID_AUDIO`. A silent recording fails.

Each instant voice counts toward the workspace's voice clone allowance. A request past it returns `400 resource_limit_reached`; deleting a voice frees its place. An instant voice cannot be retrained. To change it, create a new one.

`Idempotency-Key` is optional. Reusing a key within 24 hours returns the first response, and a duplicate that arrives while the first is still being accepted returns `409 request_in_progress`. Without the header, each retry creates a new voice.

## 2. Wait until it is ACTIVE

Poll [`GET /v3/models/audio/voices/{voice_id}`](/reference/get-an-audio-voice) until `status` is `ACTIVE` or `FAILED`.

```bash theme={null}
curl "https://api.heygen.com/v3/models/audio/voices/7c2a9e4f1b3d4e6a8f0c2d4e6a8b1234" \
  -H "X-Api-Key: $HEYGEN_API_KEY"
```

```json Response theme={null}
{
  "data": {
    "voice_id": "7c2a9e4f1b3d4e6a8f0c2d4e6a8b1234",
    "mode": "instant",
    "name": "Customer narrator",
    "language": "en",
    "status": "ACTIVE",
    "created_at": 1791565200
  }
}
```

| Status | Meaning |
| - | - |
| `PENDING` | HeyGen is checking and cleaning the recording. |
| `ACTIVE` | The voice can speak. |
| `FAILED` | Final. `failure_reason` says why. |

| `failure_reason` | Meaning |
| - | - |
| `INVALID_AUDIO` | The recording could not be fetched or read. |
| `PREPROCESSING_FAILED` | The recording did not pass the checks. Without `language`, it can also mean the language could not be detected. Send `language` to rule that out. |
| `DESIGN_FAILED` | The recording passed the checks, but no voice could be made from it. |
| `INTERNAL_ERROR` | Something failed on HeyGen's side. Try again. |

A voice created without `language` has no `language` field until HeyGen detects it. A failed voice stays in the list as `FAILED` until you delete it.

Polling is optional. Speech on a `PENDING` voice returns `409 voice_not_ready` with a `Retry-After` header; wait that many seconds and send it again.

## 3. Generate speech

Pass the voice to [Text to Speech](/docs/voices/heygen-voice-speech). [`POST /v3/models/audio/tts`](/reference/generate-model-speech) returns one 44.1 kHz WAV; [`POST /v3/models/audio/tts/stream`](/reference/stream-speech) streams audio parts over Server-Sent Events from the same body.

```bash theme={null}
curl -X POST "https://api.heygen.com/v3/models/audio/tts" \
  -H "X-Api-Key: $HEYGEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "heygen-voice-1",
    "voice_id": "7c2a9e4f1b3d4e6a8f0c2d4e6a8b1234",
    "text": "Hello from my instant voice.",
    "language": "en"
  }'
```

An instant voice takes `model`, `voice_id`, `text`, `language`, and `expressiveness_boost` (`0.0`–`1.0`, default `1.0`), plus `with_timestamps` on the stream. `seed`, `speed`, `pitch_shift`, `pitch_variance`, and `<break>` tags work with professional voices only and return `400 invalid_parameter` for an instant voice.

## Languages

`language` takes one of these codes, both when you create a voice and when you generate speech: `ar`, `be`, `bg`, `ca`, `cs`, `da`, `de`, `el`, `en`, `es`, `fa`, `fi`, `fr`, `he`, `hi`, `hr`, `hu`, `id`, `it`, `ja`, `ko`, `mk`, `ms`, `nl`, `pl`, `pt`, `ro`, `ru`, `sk`, `sl`, `sr`, `sv`, `ta`, `th`, `tl`, `tr`, `uk`, `vi`, `zh`. Any other value returns `400 invalid_parameter`.

## Manage voices

[`GET /v3/models/audio/voices`](/reference/list-audio-voices) lists the workspace's instant and professional voices, newest first; `mode` tells them apart. Set `limit` from 1 to 100 (default 10) and pass `next_token` as `token` while `has_more` is `true`.

[`DELETE /v3/models/audio/voices/{voice_id}`](/reference/delete-an-audio-voice) deletes a voice. Deleting an instant voice that is still `PENDING` cancels it.

```bash theme={null}
curl -X DELETE "https://api.heygen.com/v3/models/audio/voices/7c2a9e4f1b3d4e6a8f0c2d4e6a8b1234" \
  -H "X-Api-Key: $HEYGEN_API_KEY"
```

## Pricing

| | |
| - | - |
| Create a voice | Free during the preview. Counts toward the workspace's voice clone allowance. |
| Speech | \$15 per million characters of submitted text, spaces and punctuation included. Each request is rounded up to whole billing units of 1/60 API credit (about \$0.0167), which covers up to 1,111 characters. Only successful requests are charged; a failed or disconnected stream is free. |

| Characters in one request | Charge |
| - | - |
| 10 | \$0.0167 |
| 1,111 | \$0.0167 |
| 1,112 | \$0.0333 |
| 5,000 | \$0.0833 |

## Errors

| HTTP status | Error code | Meaning |
| - | - | - |
| `400` | `invalid_parameter` | A missing, unknown, or invalid field, more than one `audio` item, or an unsupported `language`. |
| `400` | `resource_limit_reached` | The workspace has reached its voice clone allowance. Delete a voice to free a place. |
| `403` | `resource_access_denied` | The workspace is not allowed to create voice clones. |
| `409` | `request_in_progress` | The same `Idempotency-Key` is still being accepted. Retry with backoff. |
| `429` | `rate_limit_exceeded` | Retry after the seconds in `Retry-After`. |

Speech errors are on [Text to Speech](/docs/voices/heygen-voice-speech#errors). The full catalog is in [Error Codes](/docs/error-codes).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.