
VoxCPM
About
OpenBMB's tokenizer-free Text-to-Speech system: directly generates continuous speech representations via diffusion autoregressive architecture. 2B parameter model, 30 languages, Voice Design, controllable voice cloning, 48kHz studio-quality audio.
Our Verdict
Worth TryingA tokenizer-free TTS breakthrough from OpenBMB — 30 languages, voice design, and 48kHz output, but GPU-heavy.
VoxCPM2 represents a significant advance in open-source TTS: its tokenizer-free diffusion autoregressive architecture produces more natural speech than discrete-token alternatives, and the single model covering 30 languages with automatic detection is impressive. The Voice Design feature (create a voice from a text description) and Ultimate Cloning (reproduce every nuance from reference audio + transcript) are genuinely useful capabilities.\n\nThe caveats: real-time inference needs an RTX 4090-class GPU, documentation is scattered, and the 101 open issues suggest ongoing development. For researchers, content creators, and developers who need high-quality multilingual TTS with voice design capabilities, VoxCPM2 is a compelling open-source option.
Best for
- •Researchers, content creators, and developers needing high-quality multilingual TTS with voice design and cloning, who have access to a powerful GPU (RTX 4090 or better).
Consider alternatives if
- •If you lack a powerful GPU, consider cloud-based TTS APIs (ElevenLabs, OpenAI TTS) or lighter open-source models that run on CPU.
Supported Platforms
Available platforms include Web App, Windows, macOS, Linux, and API.
Key Features
Pricing
Use Cases
Pros
Cons
Latest Update
July 2026: VoxCPM2 released with 2B parameters, 30 languages, Voice Design, controllable cloning, 48kHz output, and real-time streaming.
Related Audio & Speech Tools
Leading AI voice synthesis and cloning platform with multi-language support
OpenAI's open-source speech recognition model for multi-language speech-to-text
Voice generation platform from the team behind the open-source TTS star Fish Speech. Clone a voice from just 10-30 seconds of audio; the S1/S2 models deliver natural, expressive speech with commercial use and pay-as-you-go API
MiniMax's voice generation platform. The Speech model family delivers hyper-realistic TTS in 40+ languages with 10-second voice cloning, controllable emotion and sound-effect tags, and a free web trial