Skip to main content

Overview

The Audio API provides voice cloning and text-to-speech capabilities. Clone voices from audio samples and generate natural-sounding speech in 30+ languages. Perfect for voiceovers, audiobooks, virtual assistants, and character dialogue.

Base URL

Core Endpoints

Voice Clone Object

Audio Generation Object

Pricing

Voice Cloning

  • Cost: 5 credits per voice clone (one-time)
  • Reusable: Generate unlimited audio with same voice ID

Audio Generation

  • Base: 5 credits
  • Duration: +1 credit per second
  • Formula: Total = 5 + (duration_seconds × 1)

Examples

Quick Start

Clone a Voice

Generate Speech

Supported Languages

  • en-US - American English
  • en-GB - British English
  • en-AU - Australian English
  • en-CA - Canadian English
  • en-IN - Indian English

Preset Voices

Use built-in voices without cloning:

Voice Modulation

Control speech characteristics:

Audio Quality Options

Output Formats

Sample Rates

  • 22050 Hz: Voice-only, minimal quality
  • 44100 Hz: CD quality, recommended for most uses
  • 48000 Hz: Professional audio, broadcast quality

SSML Support

Use Speech Synthesis Markup Language for fine control:

Error Responses

Common error codes:
  • invalid_audio_sample - Poor quality or too short
  • unsupported_language - Language not available
  • text_too_long - Exceeds maximum length (10,000 chars)
  • voice_not_found - Voice ID doesn’t exist
  • insufficient_credits - Not enough credits

Rate Limits

Best Practices

For best voice cloning results:
  • Use clean, noise-free environment
  • 15-60 seconds of clear speech
  • Professional microphone recommended
  • Sample rate: 44.1kHz or higher
  • Format: WAV, MP3, or FLAC
  • Keep sentences natural and conversational
  • Use punctuation for proper pauses
  • Break long text into smaller chunks
  • Specify pronunciation for technical terms
  • Test with short samples first
  • Only clone voices you have permission to use
  • Don’t impersonate without consent
  • Clearly disclose AI-generated content
  • Comply with local voice biometric laws

Webhooks

Subscribe to audio completion events:

Next Steps

Clone Voice

Create voice profiles

Generate Speech

Text-to-speech

Add to Video

Sync with video