Text to speech that runs in your browser

Paste a script, pick one of 28 English voices, download the audio. Kokoro runs on your own machine, so the text is never uploaded and there is nothing to sign up for.

Runs on
checking…
First download
—
Voices
28 English
Model licence
Apache-2.0

0/3,000

Pick a voice

Press ▶ on any voice to hear it read one line, generated right now on your machine — these are not pre-recorded clips.

American · female

American · male

British · female

British · male

The letters are Kokoro's own quality grades from the model card: A voices were trained on the most audio, F voices on the least. Grade C and below can slur unusual words.

1.00×

Ready when you are.

How it reads your text

1 · Numbers get rewritten

Speech models read phonemes, not digits. Before anything else the text is rewritten the way a person would say it: “$4.50” becomes “4 dollars and 50 cents”, “2:30” becomes “2 30”, “1984” becomes “19 84”, “Dr.” becomes “Doctor”, and thousands separators are dropped so “12,000” is not read digit by digit.

2 · Sentences are grouped into passes

Kokoro can only see 510 phonemes at a time — roughly 500 characters of English. The text is split on sentence ends (with abbreviations like “Mr.” and decimals like “3.5” left alone), then the sentences are packed into passes that fit. Each pass starts playing as soon as it is finished, so you hear the first sentence while the rest is still being generated.

3 · Letters become phonemes, then sound

A WebAssembly build of eSpeak NG turns each sentence into IPA phonemes, a character-level tokenizer maps them to ids, and Kokoro — an 82-million-parameter StyleTTS 2 model — turns them plus a 256-number voice vector into a 24 kHz waveform. The voice vector is the only thing that differs between the 28 voices.

Measured on an M2 Max in Chrome: the “Load a chapter” text (about 1,850 characters, 111 seconds of audio) took 9.3 s on WebGPU and 166 s in CPU mode, which runs on a single WebAssembly thread; a second visit is ready in about 1.4 s from cache. Model: onnx-community/Kokoro-82M-v1.0-ONNX, fp32 for WebGPU and q4 for CPU.

Questions

Does my text get uploaded anywhere?

No. The page downloads the model from a CDN the first time, and after that everything — text normalisation, phonemes, the neural network, the WAV file — happens inside this tab. Open your browser's network panel, press Speak, and you will see no request going out.

There is also nothing on the server to store it in: this is a static page with no accounts and no database.

Which languages and voices can it read?

English only: 11 American female voices, 9 American male, 4 British female and 4 British male — 28 in total, all of them from Kokoro v1.0.

Kokoro itself ships 54 voices, including Mandarin, Japanese, Spanish, French, Italian, Portuguese and Hindi. This page does not offer them, for the reason in the next answer.

Why not the Chinese or Japanese voices?

Kokoro does not read letters, it reads phonemes, so every language needs a grapheme-to-phoneme step in front of the model. The English one (eSpeak NG compiled to WebAssembly, about 1.3 MB) runs in the browser. Mandarin needs word segmentation plus pinyin-to-IPA with tone handling, and Japanese needs a full morphological analyser and dictionary — neither exists as a browser-sized package today.

Shipping them half-working would mean a page that mangles Chinese, so the voices are simply not listed. The Mac app at the bottom of this page does Chinese TTS properly.

Can I use the audio in something commercial?

Yes. Kokoro v1.0 is released under Apache-2.0, which permits commercial use, and the full licence text is bundled with this page at LICENSE-model.txt. The text you typed stays yours — it never reached us in the first place.

The voices are synthetic and not clones of a named performer, so there is no separate voice-talent release to chase.

Why is there a 310 MB download?

Because the model runs on your machine rather than on a server, the weights have to get to your machine once. WebGPU uses the fp32 file (310 MB); CPU mode uses a 4-bit quantised file (291 MB). Both are cached by the browser, so the second visit loads in about 1.4 seconds and works with the network off.

Each voice you try adds another 0.5 MB — that is the 256-number style vector, one per voice.

How does it compare with ElevenLabs or ttsmaker?

Those run much larger models on their own servers, and at their best they sound better than Kokoro — especially on emotion and long-form pacing. What they need in return is your text on their machines, an account, and a monthly character budget.

Kokoro is 82 million parameters, which is small enough to run in a browser tab. It sounds clean and steady rather than expressive. The trade you are making here is a one-off 310 MB download instead of a monthly quota.

It mispronounced a name. Can I fix that?

Spell it the way it sounds. “Siobhan” read as written comes out wrong; “Shivawn” does not. The same trick works for acronyms you want spelled out (“S Q L” instead of “SQL”) and for surnames.

Punctuation is doing real work too: commas and dashes become short pauses, full stops become longer ones, and a blank line between paragraphs gives the longest break.

Why WAV instead of MP3?

The model produces 24 kHz audio and the WAV file is that, unencoded — about 2.8 MB per minute. Encoding MP3 in a browser means shipping a JavaScript encoder that is copyleft-licensed and roughly a third of this page's entire byte budget, and the browser's own encoder (WebCodecs) hands back raw AAC frames with no container to put them in.

So: WAV here, and the Mac app exports MP3 and M4B chapters if that is what you need.