Beginner 10 minAudio

Kokoro TTS: a convincing local voice and lightweight

Direct response

Kokoro (hexgrad/Kokoro-82M, Apache 2.0 license) is an open 82-million-parameter speech synthesis model built on StyleTTS 2 and ISTFTNet. It covers 8 languages and 54 voices (including French), runs faster than real time on a CPU, and was trained exclusively on royalty-free audio—with no cloning possible, since no cloning module has been released. Observed API cost: less than $1 per million characters.

Most good voice synthesis systems run in the cloud or require serious hardware. Kokoro is a very small open model (hexgrad/Kokoro-82M repository, Apache 2.0 license) with convincing output that generates faster than real time on an ordinary machine—even without a graphics card. Its technical documentation is unusually precise about what it contains, including the origin of its training data, making it a rare voice model whose every claim can be verified rather than simply taken on faith.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#Why a small model changes the calculation

Text-to-speech quality stopped being the problem a while ago; cost and localization are. Online voices sound very good, but they mean every sentence spoken by your system is sent to a provider and billed based on usage. Large local models sound good and occupy a graphics card you would rather allocate to your language model than to the voice that speaks for it.

A model small enough to run in one second on a processor removes both constraints at once. Reading a long document aloud is no longer a budget decision. An assistant can speak all its responses instead of only the important ones. A batch process can narrate one hundred files overnight on hardware you already own, without sending anything to a third party, waiting for a quota, or monitoring a bill that rises with the amount processed.

#The model, in verified figures

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
82 million parameters
An order of magnitude below the models it is usually compared with.
StyleTTS 2 + ISTFTNet architecture
A decoder-only design, without diffusion, which partly explains its light weight.
8 languages, 54 voices (v1.0)
Published on January 27, 2025, trained on a few hundred hours of audio.
Total training cost: about $1,000
1,000 hours of computation on an A100 80 GB GPU in total, a tiny budget for this result compared with competing models.
Data exclusively free of rights restrictions
Public-domain audio, audio under permissive licenses, or synthetic audio from closed commercial providers—explicitly never an unauthorized voice clone or audio from another open model.

This last point deserves emphasis: the model page is explicit about the provenance of the data, which is rare in this field. That is one reason Kokoro can be deployed without the legal uncertainty surrounding models trained on voices collected or cloned without permission—a point that matters particularly for a project intended for the general public in France, where voice-image rights are taken seriously. Upon release, Kokoro took first place in the community TTS Spaces Arena ranking, ahead of much larger models: XTTS v2 (467 million parameters), MetaVoice (1.2 billion), and Fish Speech (about 500 million) were all surpassed by an 82-million-parameter model—a dated ranking, not a universal guarantee, but a strong signal about the quality-to-size ratio for anyone who needs to choose quickly without running ten models in parallel.

#What it does—and does not do

Actual capabilities
CapacityKokoro
Natural prosody in ordinary proseStrong for its size
Preset voices54 voices across 8 languages, including French (fr-fr)
Voice cloning from a sampleThat isn’t its purpose: no cloning module was published with the weights
Fine-grained emotional directionLimited
SpeedFaster than real time on modest hardware
Offline operationYes, once the weights have been downloaded via pip install kokoro

The honest formulation: it's a very good reader, not an actor. For narration, notifications, an assistant, or accessibility, it's more than sufficient. For performance work where a precise intention matters, you'll need a tool that can be directed more finely. In terms of cost, the observed API rate is around $1 per million input characters—about one thousand characters per minute of output—a useful benchmark even locally for comparing against the actual long-term cost of an equivalent cloud subscription.

#Its place in a local pipeline

  1. 01
    Voice input
    A transcription model converts the user's speech into text; Faster-Whisper handles this step well, even without a GPU.
  2. 02
    The model is reasoning
    A local model produces the response, served by Ollama or another engine.
  3. 03
    The voice cuts out
    Kokoro reads the response through the kokoro Python package, which relies on the misaki phonemization library and falls back to espeak-ng for some out-of-dictionary words.

Latency must always be considered as a whole: the user perceives the total—transcription, then generation, then synthesis. Synthesis is generally the least expensive of the three steps, which means a voice assistant's perceived responsiveness is determined primarily by the speed of your language model, not by the voice that reads the final answer.

→
Stream the voice as the text is generated
Waiting for the complete response before starting to speak wastes the seconds the model spent producing it. Summarizing sentence by sentence as the text arrives makes the system dramatically more responsive without making it faster.

#Three practical uses in French

Kokoro’s fr-fr voice makes the model directly usable for a French-speaking audience, without relying on a poorly accented English voice or a third-party usage-based service. Three use cases come up most often among projects that adopt it.

Reading long documents
Reports, articles, monitoring notes: narrating a document several pages long now costs only a little CPU time, instead of a character-billed API subscription.
Home or professional voice assistant
Combined with Whisper for listening and a local LLM for reasoning, Kokoro closes the loop without any conversation leaving the local network.
Accessibility and notifications
Screen reading, voice alerts, action confirmation: use cases where latency and offline availability matter more than interpretive nuance.

In all three cases, the budget equation changes fundamentally: instead of a per-character cost charged on every call, the expense becomes the electricity and processor time of a machine you already own. For a high volume of text read aloud—from daily document monitoring to a voice customer service—the gap quickly favors local hosting, especially since throughput remains faster than real time even without a dedicated graphics card.

#Kokoro or Piper

Two different priorities
KokoroPiper
Output qualityHigher, more natural prosodyBetter, flatter
Footprint82 million parameters, a single modelVery small, one ONNX model per voice (VITS)
Processor speedFastExtremely fast, designed for highly constrained hardware
Language coverage8 languages, 54 voices (v1.0)44 languages/locales in the official voice list
LicenseApache 2.0, one license for all voicesGPL-3.0 on the active repository; MIT on the former repository, now archived
When to chooseQuality comes firstFootprint, latency, or a voice from a specific locale takes priority

Piper's licensing situation is worth checking before you commit: the historical rhasspy/piper repository, under the MIT license, is now archived and read-only. Active development has moved to OHF-Voice/piper1-gpl, maintained by the Open Home Foundation under the GPL-3.0 license — a real difference for commercial use compared with tutorials that still cite the old repository's MIT license. Piper's official documentation also notes that some individual voices may carry more restrictive licenses than the engine: check voice by voice, not just at the project level. Piper is still built differently from Kokoro: each voice is a separate VITS model exported to ONNX, with an .onnx file and an .onnx.json configuration file, whereas Kokoro provides its 54 voices from a single shared set of weights.

#Licensing and ethics

Two questions determine whether a text-to-speech model is usable in your project, and only one is technical. The first is the license for the weights and voices, which governs commercial use: for Kokoro, it is Apache 2.0, a permissive and unambiguous license covering all 54 voices as a package—this is not true of every engine, with Piper being the closest example after its recent move to GPL-3.0. The second is that a model with predefined voices completely avoids the question of consent, whereas cloning someone’s voice from a sample requires their consent and, in a growing number of jurisdictions, carries legal consequences. Kokoro’s model card is explicit: no cloning module was published with the weights, ruling out this use by design rather than through a simple usage policy.

For a French project, this combination simplifies many things: no voice licensing agreement to negotiate, no real person whose consent you need to obtain, and an Apache 2.0 license that raises no questions about resale or integration into a paid product. The only remaining point to watch is purely technical: verify that the selected French voice sounds correct on the project’s real texts, including accents and acronyms, before deploying it at scale.

#FAQ

Can Kokoro be used commercially?+
Yes. Kokoro-82M is released under the Apache 2.0 license, one of the most permissive open licenses, and its model page states that it “can be deployed anywhere, from a production environment to a personal project.” No royalties or attribution beyond the standard Apache terms are required, and this covers all 54 voices in the same block.
Do you need a graphics card?+
No. Its 82 million parameters make generation on a processor practical, generally faster than real time on ordinary hardware. A GPU helps with high-volume batch processing but is not necessary for interactive use in a voice assistant.
Can it clone my voice?+
No, and this is a deliberate design choice: the model card confirms that no cloning module was released with the weights. Instead, it provides 54 preset voices in 8 languages. Voice cloning belongs to a different class of models and raises consent questions that preset voices avoid entirely.
Kokoro or Piper?+
Kokoro for more natural output with 82 million parameters; Piper when minimal footprint and the lowest latency matter most, typically on embedded hardware. Both work entirely offline, but check the license: Kokoro is Apache 2.0, while the actively maintained Piper has moved to GPL-3.0.
Which languages are supported, and is French one of them?+
Yes: version 1.0 covers 8 languages and 54 voices, including American and British English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese, and Mandarin. Check the exact voice availability by language in the model's VOICES.md file before committing to a specific voice.
Where does Kokoro's training data come from?+
From a few hundred hours of audio, exclusively from the public domain, under a permissive license, or synthetic audio generated by closed commercial providers — explicitly never audio from another open model or an unauthorized voice clone, according to the provenance statement published on the model card.
What is the actual cost compared with a cloud API?+
Locally, the only recurring expense is electricity, for an 82-million-parameter model that runs in a few seconds on a CPU. As a benchmark, the observed market rate for Kokoro served through an API is around $1 per million characters; this is the useful comparison figure to consider before choosing between self-hosting and a third-party service.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.