Measure browser text-to-speech speed on your computer
Select one or more models and generate the same short passage with each one. The benchmark measures how long your own browser takes, making it easier to choose between a smaller CPU model and a larger WebGPU option.
Models not yet on this device download once and are cached for later runs. Large models aren't pre-selected — tick them to include their download in the benchmark. Every engine reads the same short AIVoices script using its standard narrator.
Reference results
These published and maintainer-reported results provide context for different hardware. Lower RTF is faster; 1.0 means generation took as long as the finished audio. Last reviewed July 2026.
| Model | Hardware / runtime | RTF | Source |
|---|---|---|---|
| Kokoro-82M | M2 MacBook Air (WebGPU, fp32) | 0.2–0.4 | webml-community demo, 2026 |
| Kokoro-82M | Mid-range x86 laptop (WASM q8, 1 thread) | ≈1.1 | kokoro-js community reports, 2025 |
| Piper (medium voice) | Mid-range x86 laptop (WASM) | ≈0.11 | sherpa-onnx / piper measurements, 2025 |
| KittenTTS Nano | Core i7 laptop (WASM int8) | ≈0.4 | KittenML issue tracker, 2026 |
| Kyutai Pocket TTS | Apple M4 (WASM, streaming) | ≈0.17 (6× RT) | Kyutai release notes, Jan 2026 |
| Chatterbox 0.5B | Desktop RTX-class GPU (WebGPU q4) | ≈1.0 | Resemble AI transformers.js demo, 2026 |
Methodology
- Every selected engine reads the same short script with its standard narrator, keeping the workload consistent.
- The timer starts with speech generation and ends when the audio buffer is ready. Runs that also download a model are marked with an asterisk.
- Each model uses its supported browser path, including WebGPU or WebAssembly where available.
- RTF = wall-clock generation time ÷ output audio duration.
- The results remain in the browser unless you copy them yourself.
Model download sizes and hardware requirements are on each model page: Kokoro-82M, Piper, KittenTTS Nano, Kyutai Pocket TTS, Supertonic-3, Chatterbox, MOSS-TTS-Nano, Chatterbox Multilingual, OuteTTS-1.0 (0.6B), Orpheus-3B, MOSS-TTS (8B).
Frequently asked questions
What is RTF (real-time factor)?
Real-time factor divides generation time by the length of the finished audio. An RTF of 0.5 means a ten-second clip took five seconds to generate. Values below 1.0 are faster than playback time.
Why do my numbers differ from the reference table?
Browser model speed changes with the processor, graphics hardware, available memory, browser runtime, and whether WebGPU or a CPU path was used. Reference results are context, not a promise for another device.
Is the benchmark fair between models?
Each selected model reads the same two-sentence passage with its default voice. A first run that includes a model download is marked separately. The result measures speed, not how natural or suitable the voice sounds.
Does the benchmark send my results anywhere?
No. The benchmark runs and displays its results in your browser. Use the copy button only if you choose to save or share them yourself.