Local TTS runs on your machine

Nothing you type leaves this device.

Writing for this voice

It speaks one voice

A single American English voice, cloned from Kokoro-82M's af_heart. There is no accent, gender or style control — what you get is what the model was distilled to sound like.

It was trained on plain prose

The training text was WikiText-103, so ordinary explanatory writing sits closest to what it learned. Long sentences, run-ons and fragments all flatten out. Short, well-formed sentences sound noticeably better.

Punctuation sets the pauses

The model reads your punctuation as real tokens. Sentence endings and line breaks are where a gap is inserted, and each of these is a fixed size:

. ! ?
medium
…
longer
new line
longer still
blank line
the longest break

A comma, semicolon or colon gives a brief pause of its own, but it is the model's and not a size you can choose — lengthening it would mean splitting the sentence, which sounds chopped. To make a real pause, end the sentence. An em dash (—) is read as a dash, and extra spaces do nothing.

Long sentences are handled for you

The model reads at most 510 phonemes at a time, roughly 90 words. Anything longer is split at a word boundary and stitched back together automatically, so a very long sentence still comes out whole. The only effect is a slightly longer breath where it was split. You do not need to break anything up yourself.

Numbers are read the way you would say them

Write them normally — there is no need to spell anything out. $5.50 becomes “five dollars and fifty cents”, 1st becomes “first”, 95% becomes “ninety-five percent”, 4:30 becomes “four thirty”. A bare four-digit number is read as a year, so 1995 becomes “nineteen ninety-five”; add “about” or spell it out only if you mean a quantity.

Titles and initialisms are handled too: Dr. is “doctor”, Mrs. is “missus”, and U.S.A. is spelled out letter by letter. The one thing still worth checking is proper nouns and jargon the dictionaries do not cover, which are guessed at — listen back over any name that matters.

Speed changes more than pace

1.00× is what the model was trained at, and it is the most natural. Moving away from it in either direction flattens the intonation, so treat the slider as a deliberate choice rather than a default to leave alone.

Model

Not loaded

The 35 MB model is downloaded once and then kept on this device. After that it works offline.

0 characters 0 sentences
0.60×1.60×

Punctuation sets the pauses — a comma is brief, a blank line is a full break.

Examples

Generated audio appears here.