Kyutai Pocket TTS vs Kokoro-82M
See how the two models differ in download size, language coverage, runtime needs, voice features, and output format. When both are available, use the same sentence to decide which voice works better for your project.
| Kyutai Pocket TTS | Kokoro-82M | |
|---|---|---|
| Parameters | 100M | 82M |
| Download | 189 MB | 88 MB CPU / 310 MB WebGPU |
| Speed | Begins producing audio during generation on a current multi-core CPU. | Runs around real time on a current CPU and becomes faster when WebGPU is available. |
| Voices | 28 | 28 |
| Languages | 1 | 2 |
| Voice cloning | Yes | No |
| Streaming | Yes | Yes |
| Hardware | Any CPU (WASM) | CPU ok, WebGPU faster |
| License | CC-BY-4.0 (weights), MIT (code) | Apache-2.0 |
| Output | 24.0 kHz | 24.0 kHz |
Which should you choose?
- Kyutai Pocket TTS if its main focus matches your project: streaming english speech with optional voice references.
- Kokoro-82M if its main focus matches your project: natural english speech with a practical browser download.
- Kyutai Pocket TTS if a smaller download is important; its listed download is 189 MB.
- Kyutai Pocket TTS if you want to clone a voice from reference audio.
Frequently asked questions
Which sounds better, Kyutai Pocket TTS or Kokoro-82M?
There is no universal winner. Listen for pronunciation, pacing, and tone on the script you actually plan to use, then weigh that result against download size and device requirements. Both models are available here, so you can generate the same sentence with each.
Where do the speed and size details come from?
Model sizes and runtime requirements come from the maintained model records and linked upstream sources. Use the benchmark page to measure generation speed on your own device.