Streaming and voice reference

Use Pocket TTS for streaming speech and voice references

Kyutai Pocket TTS begins returning audio while the rest of the line is still being generated. It runs on the CPU, supports English, and can condition speech on a short voice reference that you have permission to use.

Your text is processed locally by the selected model.0 / 5000

Your browser keeps this engine after its first load, so later clips can work offline. Review or remove local engines on the Models page.

Choose Pocket TTS for interactive or referenced voices

Use Pocket TTS when fast first audio or a permitted voice reference is central to the project. Its 189 MB download is larger than Piper or KittenTTS, and the quality of a referenced voice depends on the recording you provide.

Start with a clean reference and a short line

  1. 1Use only a voice recording you own or have permission to use, with one clear speaker and little background noise.
  2. 2Generate a short sentence first and check whether the identity, clarity, and pace remain stable.
  3. 3Move to longer sections only after the reference and written script produce a consistent result.

Specifications

DeveloperKyutai
Parameters100 million
Download size189 MB (int8)
Output24.0 kHz mono
Browser voices10
LanguagesEnglish
Voice cloningYes
StreamingYes — audio starts before generation finishes
LicenseCC-BY-4.0 (weights), MIT (code)
SpeedBegins producing audio during generation on a current multi-core CPU.
RuntimeWebAssembly (CPU)

Explore related downloads

The browser studio and downloadable catalog are separate. Use the catalog to find independently published model files and check each source before downloading.

Good fit when

  • +Streaming playback during generation
  • +Supports short voice-reference recordings
  • +Runs locally on a CPU

Before you choose it

  • English only in the current browser implementation
  • 189 MB initial download
  • Reference quality depends on the source recording

Voices (10)

Sleeper — Warm mature narrator · en-USBruno — Deep, heavy character voice · en-USStarlight — Bright animated character voice · en-USInferno — Low fantasy creature voice · en-USBluster — Exaggerated cartoon voice · en-USOrbit — Dry presentation voice · en-USReefy — Small upbeat cartoon voice · en-USBrass — Theatrical announcer voice · en-USVexa — Focused action character voice · en-USRiftDoc — Raspy eccentric character voice · en-US

See it alongside another model

Questions about Kyutai Pocket TTS

Does AIVoices charge to use Kyutai Pocket TTS?

AIVoices does not charge generation credits for this browser model. The listed model license is CC-BY-4.0 (weights), MIT (code). Check the upstream model and voice terms before use.

What hardware does Kyutai Pocket TTS need?

A current desktop browser and CPU are enough for this model. The 189 MB model is stored in your browser after the first download.

Is my text private when using Kyutai Pocket TTS here?

The supported model executes inside your browser tab. AIVoices does not send the text to a server for speech generation.

Can I use Kyutai Pocket TTS output commercially?

Commercial use depends on the model, selected voice, and source-data terms. This page lists CC-BY-4.0 (weights), MIT (code) for the model, but you should still review the current upstream terms.