Kokoro TTS: a convincing local voice and lightweight
Kokoro (hexgrad/Kokoro-82M, Apache 2.0 license) is an open 82-million-parameter speech synthesis model built on StyleTTS 2 and ISTFTNet. It covers 8 languages and 54 voices (including French), runs faster than real time on a CPU, and was trained exclusively on royalty-free audio—with no cloning possible, since no cloning module has been released. Observed API cost: less than $1 per million characters.
Most good voice synthesis systems run in the cloud or require serious hardware. Kokoro is a very small open model (hexgrad/Kokoro-82M repository, Apache 2.0 license) with convincing output that generates faster than real time on an ordinary machine—even without a graphics card. Its technical documentation is unusually precise about what it contains, including the origin of its training data, making it a rare voice model whose every claim can be verified rather than simply taken on faith.
#Why a small model changes the calculation
Text-to-speech quality stopped being the problem a while ago; cost and localization are. Online voices sound very good, but they mean every sentence spoken by your system is sent to a provider and billed based on usage. Large local models sound good and occupy a graphics card you would rather allocate to your language model than to the voice that speaks for it.
A model small enough to run in one second on a processor removes both constraints at once. Reading a long document aloud is no longer a budget decision. An assistant can speak all its responses instead of only the important ones. A batch process can narrate one hundred files overnight on hardware you already own, without sending anything to a third party, waiting for a quota, or monitoring a bill that rises with the amount processed.
#The model, in verified figures
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- 82 million parameters
- An order of magnitude below the models it is usually compared with.
- StyleTTS 2 + ISTFTNet architecture
- A decoder-only design, without diffusion, which partly explains its light weight.
- 8 languages, 54 voices (v1.0)
- Published on January 27, 2025, trained on a few hundred hours of audio.
- Total training cost: about $1,000
- 1,000 hours of computation on an A100 80 GB GPU in total, a tiny budget for this result compared with competing models.
- Data exclusively free of rights restrictions
- Public-domain audio, audio under permissive licenses, or synthetic audio from closed commercial providers—explicitly never an unauthorized voice clone or audio from another open model.
This last point deserves emphasis: the model page is explicit about the provenance of the data, which is rare in this field. That is one reason Kokoro can be deployed without the legal uncertainty surrounding models trained on voices collected or cloned without permission—a point that matters particularly for a project intended for the general public in France, where voice-image rights are taken seriously. Upon release, Kokoro took first place in the community TTS Spaces Arena ranking, ahead of much larger models: XTTS v2 (467 million parameters), MetaVoice (1.2 billion), and Fish Speech (about 500 million) were all surpassed by an 82-million-parameter model—a dated ranking, not a universal guarantee, but a strong signal about the quality-to-size ratio for anyone who needs to choose quickly without running ten models in parallel.
#What it does—and does not do
| Capacity | Kokoro |
|---|---|
| Natural prosody in ordinary prose | Strong for its size |
| Preset voices | 54 voices across 8 languages, including French (fr-fr) |
| Voice cloning from a sample | That isn’t its purpose: no cloning module was published with the weights |
| Fine-grained emotional direction | Limited |
| Speed | Faster than real time on modest hardware |
| Offline operation | Yes, once the weights have been downloaded via pip install kokoro |
The honest formulation: it's a very good reader, not an actor. For narration, notifications, an assistant, or accessibility, it's more than sufficient. For performance work where a precise intention matters, you'll need a tool that can be directed more finely. In terms of cost, the observed API rate is around $1 per million input characters—about one thousand characters per minute of output—a useful benchmark even locally for comparing against the actual long-term cost of an equivalent cloud subscription.
#Its place in a local pipeline
- 01Voice inputA transcription model converts the user's speech into text; Faster-Whisper handles this step well, even without a GPU.
- 02The model is reasoningA local model produces the response, served by Ollama or another engine.
- 03The voice cuts outKokoro reads the response through the kokoro Python package, which relies on the misaki phonemization library and falls back to espeak-ng for some out-of-dictionary words.
Latency must always be considered as a whole: the user perceives the total—transcription, then generation, then synthesis. Synthesis is generally the least expensive of the three steps, which means a voice assistant's perceived responsiveness is determined primarily by the speed of your language model, not by the voice that reads the final answer.
#Three practical uses in French
Kokoro’s fr-fr voice makes the model directly usable for a French-speaking audience, without relying on a poorly accented English voice or a third-party usage-based service. Three use cases come up most often among projects that adopt it.
- Reading long documents
- Reports, articles, monitoring notes: narrating a document several pages long now costs only a little CPU time, instead of a character-billed API subscription.
- Home or professional voice assistant
- Combined with Whisper for listening and a local LLM for reasoning, Kokoro closes the loop without any conversation leaving the local network.
- Accessibility and notifications
- Screen reading, voice alerts, action confirmation: use cases where latency and offline availability matter more than interpretive nuance.
In all three cases, the budget equation changes fundamentally: instead of a per-character cost charged on every call, the expense becomes the electricity and processor time of a machine you already own. For a high volume of text read aloud—from daily document monitoring to a voice customer service—the gap quickly favors local hosting, especially since throughput remains faster than real time even without a dedicated graphics card.
#Kokoro or Piper
| Kokoro | Piper | |
|---|---|---|
| Output quality | Higher, more natural prosody | Better, flatter |
| Footprint | 82 million parameters, a single model | Very small, one ONNX model per voice (VITS) |
| Processor speed | Fast | Extremely fast, designed for highly constrained hardware |
| Language coverage | 8 languages, 54 voices (v1.0) | 44 languages/locales in the official voice list |
| License | Apache 2.0, one license for all voices | GPL-3.0 on the active repository; MIT on the former repository, now archived |
| When to choose | Quality comes first | Footprint, latency, or a voice from a specific locale takes priority |
Piper's licensing situation is worth checking before you commit: the historical rhasspy/piper repository, under the MIT license, is now archived and read-only. Active development has moved to OHF-Voice/piper1-gpl, maintained by the Open Home Foundation under the GPL-3.0 license — a real difference for commercial use compared with tutorials that still cite the old repository's MIT license. Piper's official documentation also notes that some individual voices may carry more restrictive licenses than the engine: check voice by voice, not just at the project level. Piper is still built differently from Kokoro: each voice is a separate VITS model exported to ONNX, with an .onnx file and an .onnx.json configuration file, whereas Kokoro provides its 54 voices from a single shared set of weights.
- Local voice assistant: Whisper + Ollama + Piper end to end
- The listening stage: transcribe quickly locally, even without a GPU
- Vosk: another option for offline speech recognition
- Automate a meeting summary locally
- Source: official Kokoro-82M page on Hugging Face
- Source: the GitHub repository for the kokoro Python package
- Source: active Piper repository (GPL-3.0)
#Licensing and ethics
Two questions determine whether a text-to-speech model is usable in your project, and only one is technical. The first is the license for the weights and voices, which governs commercial use: for Kokoro, it is Apache 2.0, a permissive and unambiguous license covering all 54 voices as a package—this is not true of every engine, with Piper being the closest example after its recent move to GPL-3.0. The second is that a model with predefined voices completely avoids the question of consent, whereas cloning someone’s voice from a sample requires their consent and, in a growing number of jurisdictions, carries legal consequences. Kokoro’s model card is explicit: no cloning module was published with the weights, ruling out this use by design rather than through a simple usage policy.
For a French project, this combination simplifies many things: no voice licensing agreement to negotiate, no real person whose consent you need to obtain, and an Apache 2.0 license that raises no questions about resale or integration into a paid product. The only remaining point to watch is purely technical: verify that the selected French voice sounds correct on the project’s real texts, including accents and acronyms, before deploying it at scale.
#FAQ
Can Kokoro be used commercially?+
Do you need a graphics card?+
Can it clone my voice?+
Kokoro or Piper?+
Which languages are supported, and is French one of them?+
Where does Kokoro's training data come from?+
What is the actual cost compared with a cloud API?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.