Voice cloning from a 30-second sample.

Record or upload two to thirty seconds of real speech. FuturVoice transcribes it for you, builds the voice, and generates in it - in ten languages, from a browser, with nothing to install.

What a good sample looks like

Two to thirty seconds of ordinary speech, recorded anywhere quiet enough to hear the words. There is nothing to type: the sample is transcribed automatically, and that transcript is what the engine aligns against.

Near-silent clips are rejected outright rather than accepted and turned into a voice that sounds nothing like anyone. A rejected upload costs nothing and tells you why.

The numbers

Sample length
2–30 seconds
Sample size limit
50 MB
Languages for cloned voices
10
Cost to generate
1 credit per second of audio
Free on signup
500 credits (≈ 8 minutes)

Three engines to choose between

Voice, engine and language all sit in one box, so you can repeat a take exactly or ask for a different one. Generation is billed by the second of audio actually delivered - a failure costs nothing, and asking for exactly the same thing twice costs nothing the second time.

What happens after the voice exists

It becomes something you can build with. Takes land in a timeline you can trim, reorder and layer, eight studio effects can be tried without touching the original, and everything is exportable as WAV, MP3 or Opus.

Questions about cloning

How much audio do you need to clone a voice?

Two to thirty seconds of real speech - near-silent clips are rejected rather than turned into a bad voice. We transcribe the sample, so there is nothing to type.

Can I use it commercially?

Yes, for voices you have the right to use. Cloning someone’s voice without their permission is not something we will help with. See the terms.

Clone a voice with 500 free credits