1 · Numbers get rewritten
Speech models read phonemes, not digits. Before anything else the text is rewritten the way a person would say it: “$4.50” becomes “4 dollars and 50 cents”, “2:30” becomes “2 30”, “1984” becomes “19 84”, “Dr.” becomes “Doctor”, and thousands separators are dropped so “12,000” is not read digit by digit.
2 · Sentences are grouped into passes
Kokoro can only see 510 phonemes at a time — roughly 500 characters of English. The text is split on sentence ends (with abbreviations like “Mr.” and decimals like “3.5” left alone), then the sentences are packed into passes that fit. Each pass starts playing as soon as it is finished, so you hear the first sentence while the rest is still being generated.
3 · Letters become phonemes, then sound
A WebAssembly build of eSpeak NG turns each sentence into IPA phonemes, a character-level tokenizer maps them to ids, and Kokoro — an 82-million-parameter StyleTTS 2 model — turns them plus a 256-number voice vector into a 24 kHz waveform. The voice vector is the only thing that differs between the 28 voices.