Multilingual WebGPU

What Supertonic-3 needs to run

Supertonic-3 is a 99-million-parameter multilingual model with first-party browser support and 31 listed languages. It is aimed at projects that need one model for several languages and generates 44.1 kHz audio.

Supertonic-3 is not ready in the generator yet. Review its size, runtime needs, and current limitations below, or use one of the available models . The compatibility check can also show whether your browser exposes the hardware features this model would need.

Specifications

DeveloperSupertone
Parameters99 million
Download size398 MB (fp32)
Output44.1 kHz mono
Listed voices10
LanguagesEnglish (US), Korean, Japanese, Chinese, Spanish, French, German, Italian, Portuguese, Hindi, Arabic, Russian, Vietnamese, Indonesian, Thai, Turkish, Polish, Dutch
Voice cloningNo
StreamingNo
LicenseOpenRAIL-M (weights), MIT (code) — OpenRAIL-M allows commercial use subject to its use-based restrictions. Keep the model notice with redistributed copies.
SpeedWebGPU is recommended; CPU-only generation may be noticeably slower.
RuntimeWebAssembly, with WebGPU acceleration when available

Explore related downloads

The browser studio and downloadable catalog are separate. Use the catalog to find independently published model files and check each source before downloading.

Good fit when

  • +31 languages in one model
  • +Official browser-oriented ONNX support
  • +44.1 kHz generated audio

Before you choose it

  • 398 MB model download
  • WebGPU is recommended for practical generation speed
  • OpenRAIL-M terms include use-based restrictions

Questions about Supertonic-3

Does AIVoices charge to use Supertonic-3?

AIVoices does not charge generation credits for this browser model. The listed model license is OpenRAIL-M (weights), MIT (code). OpenRAIL-M allows commercial use subject to its use-based restrictions. Keep the model notice with redistributed copies.

What hardware does Supertonic-3 need?

The CPU path works in a current desktop browser, while WebGPU can improve generation speed. About 8 GB of RAM is recommended.

Is my text private when using Supertonic-3 here?

The supported model executes inside your browser tab. AIVoices does not send the text to a server for speech generation.

Can I use Supertonic-3 output commercially?

Commercial use depends on the model, selected voice, and source-data terms. This page lists OpenRAIL-M (weights), MIT (code) for the model and notes that OpenRAIL-M allows commercial use subject to its use-based restrictions. Keep the model notice with redistributed copies.