Natural English
Use Kokoro-82M for English text to speech
Kokoro is the strongest starting point here when natural English narration matters more than having the smallest download. Choose from 28 US and UK voices, run the model on CPU or with WebGPU acceleration, and export the finished speech as a WAV.
Your browser keeps this engine after its first load, so later clips can work offline. Review or remove local engines on the Models page.
Choose Kokoro for polished English narration
Use Kokoro for YouTube voice-overs, audiobooks, explainers, and other English scripts where pacing and delivery matter. Choose Piper for broader language support, KittenTTS for a smaller first download, or Pocket TTS when streaming and voice-reference controls matter more.
Test Kokoro with the hardest part of the script
- 1Paste a paragraph containing the names, numbers, and abbreviations used in the project.
- 2Try two or three voices and listen for pacing across both short and long sentences.
- 3Correct the written line where necessary, then generate approved sections as separate WAV files.
Specifications
| Developer | hexgrad |
|---|---|
| Parameters | 82 million |
| Download size | 88 MB CPU / 310 MB WebGPU (q8 (WASM) / fp32 (WebGPU)) |
| Output | 24.0 kHz mono |
| Browser voices | 28 |
| Languages | English (US), English (UK) |
| Voice cloning | No |
| Streaming | Yes — audio starts before generation finishes |
| License | Apache-2.0 — The model weights are Apache-2.0. The browser phonemizer includes components derived from espeak-ng, which uses GPL-3.0. |
| Speed | Runs around real time on a current CPU and becomes faster when WebGPU is available. |
| Runtime | WebAssembly, with WebGPU acceleration when available |
Specifications checked 2026-07-21 against:
Explore related downloads
The browser studio and downloadable catalog are separate. Use the catalog to find independently published model files and check each source before downloading.
Good fit when
- +Natural pacing for a compact browser model
- +Multiple US and UK English voices
- +CPU and WebGPU execution paths
Before you choose it
- −Focused mainly on English
- −Does not create a new voice from a reference recording
- −The WebGPU model is larger than the CPU version
Voices (28)
See it alongside another model
Questions about Kokoro-82M
Does AIVoices charge to use Kokoro-82M?
AIVoices does not charge generation credits for this browser model. The listed model license is Apache-2.0. The model weights are Apache-2.0. The browser phonemizer includes components derived from espeak-ng, which uses GPL-3.0.
What hardware does Kokoro-82M need?
The CPU path works in a current desktop browser, while WebGPU can improve generation speed. About 4 GB of RAM is recommended.
Is my text private when using Kokoro-82M here?
The supported model executes inside your browser tab. AIVoices does not send the text to a server for speech generation.
Can I use Kokoro-82M output commercially?
Commercial use depends on the model, selected voice, and source-data terms. This page lists Apache-2.0 for the model and notes that The model weights are Apache-2.0. The browser phonemizer includes components derived from espeak-ng, which uses GPL-3.0.