TrendWhat rose yesterday, every day at 07:30 KST · 한국어

GitHub · as of October 8, 2026 · View on GitHub →

supertone-oss-archive/supertonic

Multilingual on-device text-to-speech engine

Run real-time text-to-speech in 31 languages locally without the cloud

When you read a webpage or listen to an e-book, natural-sounding voice reaches your ears even without an internet connection. This is because the text-to-speech system runs directly inside your computer or smartphone. Supertonic is designed to keep the model resident on the device and perform inference using ONNX Runtime, providing an environment where data does not leave for external servers.

Archive status and latest version

This repository is currently in an archived state, with development and official support ended. The code and models are preserved for reference under the supertone-oss-archive organization, and no further updates or security patches are provided. The latest version, Supertonic 3, supports 31 languages and reduces repetition or omission errors compared to previous versions while improving pronunciation accuracy. Model weights can be downloaded from the Hugging Face archive namespace, and it maintains a v2-compatible interface so that previous integration code can be used as is.

Why it runs on-device

Supertonic is a small model with 99M parameters, making it much lighter than other open-source TTS systems of 0.7B to 2B scale. This allows for real-time synthesis using only the CPU without a GPU, even on resource-constrained hardware like Raspberry Pi or e-book readers. The output audio is a 44.1kHz 16-bit WAV file, maintaining quality that can be played back immediately without a separate upsampler. Due to the lack of network dependency, it is suitable for privacy-focused environments or offline work.

Natural sentence handling and expression tags

It is trained to correctly pronounce complex text such as financial notations, phone numbers, and technical units without pre-processing. For example, notations like '$5.2M' or '(212) 555-0142 ext. 402' are read appropriately according to context. Additionally, you can add human nuances like laughter or breathing sounds using 10 inline tags such as <laugh>, <breath>, and <sigh>. If you pass lang="na" without specifying a language, it automatically detects and processes the language of the input text.

Support for various language ecosystems

Sample code is provided for use in major languages and platforms such as Python, Node.js, Java, C++, C#, Go, Swift, Rust, and Flutter. In the browser, inference can be performed on the client side using WebGPU. Each language directory has an independent README where you can check installation and execution methods for that environment. The Python SDK can be installed with pip install supertonic, but to use it with archived models, you must specify a local assets/ directory.

Things to check before use

Since this project is no longer maintained, you should not expect new bug fixes or feature additions. The model is distributed under the OpenRAIL-M license, and sample code under the MIT license; you must review the license terms before commercial use. Hosting demos or the Voice Builder service are not included in the archive, and access to these services will become unavailable after August 31, 2026. Downloading and running the model in a local environment requires a network connection, but the internet is not needed during the inference stage.

By the numbers

Language
Swift
Topics
cpp · csharp · flutter · go · ios · java · lightweight · multilingual
Latest release
v2.0.0 · January 6, 2026
Last commit
September 9, 2026
Open issues
98
Open pull requests
33

Related repositories

Written by AI from this repository's README on October 7, 2026. GitHub's original is the reference.

View on GitHub → · Homepage