Large WebGPU model
What Chatterbox needs to run
Chatterbox is a 500-million-parameter English model from Resemble AI. It supports expressive speech and can condition generation on a short reference recording, but the browser download and memory requirements make it a desktop-GPU option rather than a general default.
Chatterbox is not ready in the generator yet. Review its size, runtime needs, and current limitations below, or use one of the available models . The compatibility check can also show whether your browser exposes the hardware features this model would need.
Specifications
| Developer | Resemble AI |
|---|---|
| Parameters | 500 million |
| Download size | 1550 MB (q4 mixed) |
| Output | 24.0 kHz mono |
| Listed voices | 1 |
| Languages | English (US) |
| Voice cloning | Yes |
| Streaming | No |
| License | MIT |
| Speed | Intended for desktop WebGPU hardware; CPU-only generation is not practical. |
| Runtime | WebGPU required |
Specifications checked 2026-07-21 against:
Explore related downloads
The browser studio and downloadable catalog are separate. Use the catalog to find independently published model files and check each source before downloading.
Good fit when
- +Expressive English output
- +Supports a short voice-reference recording
- +MIT-licensed model release
Before you choose it
- −Requires WebGPU and substantial memory
- −Approximately 1.5 GB initial download
Questions about Chatterbox
Does AIVoices charge to use Chatterbox?
AIVoices does not charge generation credits for this browser model. The listed model license is MIT. Check the upstream model and voice terms before use.
What hardware does Chatterbox need?
Use a current desktop browser with WebGPU, a compatible GPU, and about 8 GB of RAM. Run the compatibility check before downloading the model.
Is my text private when using Chatterbox here?
The supported model executes inside your browser tab. AIVoices does not send the text to a server for speech generation.
Can I use Chatterbox output commercially?
Commercial use depends on the model, selected voice, and source-data terms. This page lists MIT for the model, but you should still review the current upstream terms.