VoxCPM

VoxCPM

Audio & Speech
OpenBMB (Tsinghua University)
Open SourceFree

About

OpenBMB's tokenizer-free Text-to-Speech system: directly generates continuous speech representations via diffusion autoregressive architecture. 2B parameter model, 30 languages, Voice Design, controllable voice cloning, 48kHz studio-quality audio.

Share this tool

Our Verdict

Worth Trying

A tokenizer-free TTS breakthrough from OpenBMB — 30 languages, voice design, and 48kHz output, but GPU-heavy.

VoxCPM2 represents a significant advance in open-source TTS: its tokenizer-free diffusion autoregressive architecture produces more natural speech than discrete-token alternatives, and the single model covering 30 languages with automatic detection is impressive. The Voice Design feature (create a voice from a text description) and Ultimate Cloning (reproduce every nuance from reference audio + transcript) are genuinely useful capabilities.\n\nThe caveats: real-time inference needs an RTX 4090-class GPU, documentation is scattered, and the 101 open issues suggest ongoing development. For researchers, content creators, and developers who need high-quality multilingual TTS with voice design capabilities, VoxCPM2 is a compelling open-source option.

Best for

  • Researchers, content creators, and developers needing high-quality multilingual TTS with voice design and cloning, who have access to a powerful GPU (RTX 4090 or better).

Consider alternatives if

  • If you lack a powerful GPU, consider cloud-based TTS APIs (ElevenLabs, OpenAI TTS) or lighter open-source models that run on CPU.

Supported Platforms

Web AppWindowsmacOSLinuxAPI

Available platforms include Web App, Windows, macOS, Linux, and API.

Key Features

Tokenizer-free architecture: directly generates continuous speech representations via diffusion autoregressive model
2 billion parameter model (VoxCPM2) trained on 2M+ hours of multilingual speech data
30 languages supported — no language tag needed, auto-detects from input text
Voice Design: create a brand-new voice from natural-language description (gender, age, tone, emotion, pace)
Controllable Voice Cloning: clone any voice from a short reference clip with style guidance
Ultimate Cloning: reproduce every vocal nuance from reference audio + transcript
48kHz studio-quality audio output via AudioVAE V2 with built-in super-resolution
Real-time streaming: RTF as low as ~0.3 on RTX 4090, ~0.13 with Nano-vLLM acceleration
Context-aware synthesis: automatically infers prosody and expressiveness from text content
Built on MiniCPM-4 backbone from OpenBMB

Pricing

free
Free and open source (Apache 2.0). Model weights available on Hugging Face and ModelScope.

Use Cases

Multilingual TTS for content creation, dubbing, and accessibility
Voice design for virtual assistants, gaming characters, and brand voices
High-fidelity voice cloning for personalized speech applications

Pros

Tokenizer-free architecture produces more natural, expressive speech than discrete-token TTS
30 languages in one model with automatic language detection
Voice Design from natural language description — no reference audio needed
48kHz studio-quality output with built-in super-resolution
Apache 2.0 licensed, from the reputable OpenBMB team (Tsinghua University)

Cons

Requires significant GPU resources (RTX 4090 recommended) for real-time inference
101 open issues at verification time
Documentation is scattered across GitHub, ReadTheDocs, and Hugging Face
Last push was 35 days ago — moderate activity level

Latest Update

July 2026: VoxCPM2 released with 2B parameters, 30 languages, Voice Design, controllable cloning, 48kHz output, and real-time streaming.

Subscribe to AI Updates

Get the latest AI tool recommendations, industry insights, and analysis delivered to your inbox.

We respect your privacy. Unsubscribe at any time.