Home AI Tools Blogs AI News About Us Contact Us
➕ Submit AI Tools ✍️ Write for Us
Home AI Tools AI Voice Cloning VoxCPM2

VoxCPM2

✓ Verified
(4.6) Free
Web API

VoxCPM2 is an open-source voice generation model for voice design, controllable voice cloning and high-quality multilingual speech.

Best For: Developers and creators needing open-source voice cloning and local speech generation

The Verdict: VoxCPM2

VoxCPM2 is a strong choice for developers who want an open-source voice cloning model without being tied to a monthly SaaS subscription. Its 30-language support, 48kHz output, voice design and two cloning approaches give it considerably more flexibility than a basic local TTS model.

The main drawback is the setup. This isn't a simple website where you upload audio, type a sentence and forget about the technical side. Local inference requires suitable hardware, and the project's own documentation says controllable cloning and voice design can vary between runs. For developers who value local control and open licensing, though, that trade-off is reasonable.

What is VoxCPM2?

VoxCPM2 is an open-source speech model that creates multilingual speech, designs new voices and clones existing voices from reference audio. The latest 2B model supports 30 languages, voice design, controllable cloning and 48kHz audio output.

The model gives you three distinct ways to create speech. Voice Design creates a new voice from a text description, Controllable Cloning uses reference audio while allowing control over style and emotion, and Ultimate Cloning uses the reference recording and its transcript to preserve more of the speaker's vocal characteristics.

Unlike a typical voice-generation subscription, VoxCPM2 can be downloaded and run locally. The official project is released under Apache 2.0, and its current documentation reports real-time streaming performance on suitable NVIDIA hardware.

Who is VoxCPM2 Best For?

Developers and creators needing open-source voice cloning and local speech generation

VoxCPM2 Key Features

AI text to speech
Controllable voice cloning
Ultimate voice cloning
Voice design from text descriptions
Zero-shot voice creation
Reference audio voice cloning
Transcript-guided voice cloning
Emotion control
Speaking rate control
Tone and style control
Pitch variation control
Context-aware speech synthesis
Multilingual speech generation
30 supported languages
48kHz audio output
Natural breathing and prosody
Real-time streaming
Local model deployment
Apache 2.0 license
Open-source model weights
Open-source inference code
LoRA fine-tuning support
OpenAI-compatible API through supported serving setups
NVIDIA GPU acceleration
Voice cloning without model retraining

VoxCPM2 Review Summary

Performance ScoreA
Content QualityVoxCPM2 combines open-source speech synthesis, voice design, controllable cloning and multilingual generation in one model.
InterfaceThe web demo provides a simple way to test voice design and cloning, while local deployment requires a more technical setup.
AI TechnologyVoxCPM2 uses a tokenizer-free diffusion autoregressive architecture with AudioVAE V2 for expressive multilingual speech generation.
PurposeVoxCPM2 is designed for developers and researchers who want high-quality speech generation and voice cloning with local control.
CompatibilityWeb, API
Pricing SummaryVoxCPM2 is free and open source under the Apache 2.0 license, with no required subscription, while self-hosting can create hardware or cloud-computing costs.

VoxCPM2 Pros & Cons

✅ Pros

  • The model is available under the Apache 2.0 license
  • Voice cloning works from short reference audio
  • Voice Design does not require reference audio
  • Supports 30 languages
  • Produces native 48kHz audio output
  • Offers separate controllable and transcript-guided cloning modes
  • Can run locally instead of requiring a hosted subscription
  • Supports real-time streaming on suitable hardware
  • Model weights and code are publicly available
  • Supports LoRA fine-tuning for customization

❌ Cons

  • It is primarily a developer and self-hosted model rather than a polished SaaS product
  • Running the model locally requires compatible hardware
  • Voice Design and controllable cloning can vary between generations
  • The quality and speed depend heavily on the hardware and deployment setup
  • There is no standard consumer subscription or official hosted pricing plan
  • The official project requires technical setup for local use

FAQs

What is VoxCPM2 used for?

VoxCPM2 is used for text to speech, voice cloning, voice design, multilingual speech generation, audiobook narration, voiceovers, games, virtual assistants and other speech applications.

Is VoxCPM2 free to use?

Yes. The official VoxCPM2 model and its code are released under the Apache 2.0 license. There is no required subscription for downloading and running the open-source model, although local operation can require suitable hardware.

What platforms does VoxCPM2 support?

VoxCPM2 can be used through its web demo and local developer environments. The official project also supports API serving through compatible inference setups.

Who is VoxCPM2 best for?

VoxCPM2 is best for developers, researchers and technical creators who want local voice cloning, multilingual speech generation and control over their voice-generation setup.

Does VoxCPM2 offer a free trial?

VoxCPM2 does not use a conventional free trial because the official model is open source. Users can access the model and code directly and run it on compatible hardware.

How much does VoxCPM2 cost?

The official VoxCPM2 project is free under the Apache 2.0 license. There is no official monthly subscription price for the open-source model. Running it yourself may still involve hardware or cloud-computing costs.

Can VoxCPM2 clone a voice?

Yes. VoxCPM2 supports controllable voice cloning from reference audio and an Ultimate Cloning mode that uses reference audio together with its transcript to preserve more of the speaker's vocal characteristics.

How much audio is needed for VoxCPM2 voice cloning?

The official project supports short reference clips for controllable cloning. Its documentation also describes transcript-guided Ultimate Cloning for cases where users want to preserve more vocal detail from a longer reference recording.

Does VoxCPM2 support voice design?

Yes. Voice Design can create a new synthetic voice from a natural-language description without requiring reference audio. You can describe characteristics such as age, gender, tone, emotion and speaking rate.

How many languages does VoxCPM2 support?

The official VoxCPM2 release supports 30 languages, including English, Chinese, Japanese, Korean, Arabic, Hindi, French, German, Spanish, Portuguese, Italian and several other languages.

Does VoxCPM2 support real-time speech?

Yes. The official project reports real-time streaming performance with an RTF of about 0.30 on an NVIDIA RTX 4090, with faster results possible through Nano-VLLM or vLLM-based serving.

Does VoxCPM2 have an API?

VoxCPM2 can be served through an OpenAI-compatible API using supported vLLM or Nano-VLLM deployments. The official project also provides local inference options for developers.

What audio quality does VoxCPM2 support?

VoxCPM2 supports native 48kHz audio output. The official architecture uses AudioVAE V2 with built-in super-resolution to produce high-resolution speech.

Can VoxCPM2 run locally?

Yes. VoxCPM2 is an open-source model that can be downloaded and run locally. The official project lists around 8GB of VRAM for the 2B model, although actual requirements can vary with the inference setup.

Is VoxCPM2 commercially usable?

The official VoxCPM2 code and weights are released under the Apache 2.0 license, which permits commercial use subject to the license terms. Users must still have the appropriate rights to any reference voice they clone and follow applicable laws.
Scroll to Top