Skip to main content

Overview

Percify’s Voice Cloning technology enables you to create realistic voice replicas and generate natural-sounding speech in multiple languages. Perfect for adding personality to avatars, creating audio content, or building voice-enabled applications.

Natural Speech

Human-like intonation and emotion

Multi-Language

Support for 30+ languages

Quick Cloning

Clone voices from 15s samples

Flexible Output

Control speed, pitch, and emotion

How Voice Cloning Works

1

Upload Voice Sample

Provide a clean audio recording (15-60 seconds recommended)
2

AI Analysis

Our models analyze vocal characteristics, tone, and patterns
3

Voice Profile Creation

Generate a unique voice ID you can reuse
4

Text-to-Speech Generation

Convert any text to speech using the cloned voice

Voice Cloning Requirements

Audio Sample Quality

Optimal Conditions:
  • Clear, noise-free environment
  • Professional or high-quality microphone
  • Consistent volume level
  • No background music or effects
  • Sample rate: 44.1kHz or higher
  • Format: WAV, MP3, or FLAC
  • Minimum: 15 seconds (basic cloning)
  • Recommended: 30-60 seconds (better quality)
  • Maximum: 5 minutes (professional cloning)
Longer samples provide better voice fidelity and natural intonation.
Read clear, varied sentences that include:
  • Different emotions (neutral, happy, serious)
  • Various pitch ranges
  • Natural pauses and breathing
  • Complete sentences with proper intonation
Example script: “Hello, I’m excited to try voice cloning with Percify. The technology is amazing and opens up so many creative possibilities. I can imagine using this for podcasts, videos, or even virtual assistants. Let’s see how well it captures my unique voice characteristics.”

Pricing

Voice Cloning (One-time per voice)

  • Cost: 5 credits per voice clone
  • Includes: Voice profile creation and storage
  • Reusable: Generate unlimited audio with the same voice ID

Audio Generation (Per use)

  • Base Cost: 5 credits
  • Duration Cost: +1 credit per second of generated audio
  • Formula: Total = 5 + (duration_seconds × 1)

Examples

Cost Optimization: Clone a voice once (5 credits), then reuse it indefinitely. Only pay for generated audio duration.

Creating a Voice Clone

Via Dashboard

1

Navigate to Voice Studio

Click “Voice Cloning” from the main menu
2

Upload Audio Sample

Drag and drop or select your audio file (15s-5min)
3

Name Your Voice

Give your voice clone a memorable name
4

Set Language

Select primary language for the voice
5

Create Clone

Click “Clone Voice” and wait 30-60 seconds for processing

Via API

Response:

Generating Speech

Basic Text-to-Speech

Advanced Options

Supported Languages

  • en-US - American English
  • en-GB - British English
  • en-AU - Australian English
  • en-CA - Canadian English
  • en-IN - Indian English
Language detection is automatic based on input text. Specify language explicitly for best results with multilingual content.

Preset Voices

Don’t have a voice sample? Use our preset voices:

Professional Male

Deep, authoritative, news anchor style

Professional Female

Clear, confident, corporate presenter

Friendly Male

Warm, approachable, conversational

Friendly Female

Upbeat, energetic, engaging

Narrator

Storytelling, documentary style

Character Voices

Various character archetypes
Access preset voices:

Combining Voice with Video

Create fully animated, voiced avatar videos:

Voice Management

List Your Voices

Update Voice Metadata

Delete Voice

Audio Quality Optimization

  • MP3: Best for web/streaming (smaller files)
  • WAV: Highest quality, uncompressed (large files)
  • OGG: Good compression, open format
Choose based on your use case and delivery method.
Higher values = better quality but larger filesRecommended settings:
  • Web/mobile: 44.1kHz, 128kbps MP3
  • Professional: 48kHz, 320kbps MP3 or WAV
  • Podcast: 44.1kHz, 192kbps MP3
Percify automatically applies:
  • Noise reduction
  • Volume normalization
  • De-essing (reduces harsh ‘s’ sounds)
  • Breath removal (optional)
Disable with applyProcessing: false for raw output.

Advanced Features

SSML Support

Use Speech Synthesis Markup Language for fine control:

Phoneme-Level Control

Specify exact pronunciation:

Voice Mixing

Combine multiple voices in one audio:

Use Cases

Content Creation

YouTube videos, podcasts, audiobooks

E-Learning

Online courses, training materials

Accessibility

Screen readers, audio descriptions

Gaming

Character dialogue, narration

Virtual Assistants

Chatbots, voice interfaces

Marketing

Ads, promotional content

Best Practices

Recording Tips:
  • Use a pop filter to reduce plosives (p, b sounds)
  • Record in a quiet, carpeted room
  • Maintain consistent distance from mic (6-12 inches)
  • Speak naturally, don’t over-enunciate
  • Do multiple takes and choose the best
Content Policy:
  • Only clone voices you have permission to use
  • Don’t impersonate others without consent
  • No deceptive or fraudulent use
  • Comply with local voice biometric laws
  • Clearly disclose AI-generated content when required

Troubleshooting

API Rate Limits

Next Steps

Image to Video

Combine voice with animation

Audio API Reference

Complete API docs

Credits System

Voice pricing details

Support

Need help with voice cloning? Visit the FAQ or email support@percify.io.