Measure browser text-to-speech speed on your computer

Select one or more models and generate the same short passage with each one. The benchmark measures how long your own browser takes, making it easier to choose between a smaller CPU model and a larger WebGPU option.

Models not yet on this device download once and are cached for later runs. Large models aren't pre-selected — tick them to include their download in the benchmark. Every engine reads the same short AIVoices script using its standard narrator.

Reference results

These published and maintainer-reported results provide context for different hardware. Lower RTF is faster; 1.0 means generation took as long as the finished audio. Last reviewed July 2026.

ModelHardware / runtimeRTFSource
Kokoro-82MM2 MacBook Air (WebGPU, fp32)0.2–0.4webml-community demo, 2026
Kokoro-82MMid-range x86 laptop (WASM q8, 1 thread)≈1.1kokoro-js community reports, 2025
Piper (medium voice)Mid-range x86 laptop (WASM)≈0.11sherpa-onnx / piper measurements, 2025
KittenTTS NanoCore i7 laptop (WASM int8)≈0.4KittenML issue tracker, 2026
Kyutai Pocket TTSApple M4 (WASM, streaming)≈0.17 (6× RT)Kyutai release notes, Jan 2026
Chatterbox 0.5BDesktop RTX-class GPU (WebGPU q4)≈1.0Resemble AI transformers.js demo, 2026

Methodology

  • Every selected engine reads the same short script with its standard narrator, keeping the workload consistent.
  • The timer starts with speech generation and ends when the audio buffer is ready. Runs that also download a model are marked with an asterisk.
  • Each model uses its supported browser path, including WebGPU or WebAssembly where available.
  • RTF = wall-clock generation time ÷ output audio duration.
  • The results remain in the browser unless you copy them yourself.

Model download sizes and hardware requirements are on each model page: Kokoro-82M, Piper, KittenTTS Nano, Kyutai Pocket TTS, Supertonic-3, Chatterbox, MOSS-TTS-Nano, Chatterbox Multilingual, OuteTTS-1.0 (0.6B), Orpheus-3B, MOSS-TTS (8B).

Frequently asked questions

What is RTF (real-time factor)?

Real-time factor divides generation time by the length of the finished audio. An RTF of 0.5 means a ten-second clip took five seconds to generate. Values below 1.0 are faster than playback time.

Why do my numbers differ from the reference table?

Browser model speed changes with the processor, graphics hardware, available memory, browser runtime, and whether WebGPU or a CPU path was used. Reference results are context, not a promise for another device.

Is the benchmark fair between models?

Each selected model reads the same two-sentence passage with its default voice. A first run that includes a model download is marked separately. The result measures speed, not how natural or suitable the voice sounds.

Does the benchmark send my results anywhere?

No. The benchmark runs and displays its results in your browser. Use the copy button only if you choose to save or share them yourself.