Free Text to Speech — Download as MP3 or WAV

Paste your script, pick a voice, and get a natural-sounding recording you can download as an MP3 or a WAV. There is no sign-up, no watermark and no character-per-month meter, because the model generating the voice runs inside this page rather than on a server somewhere charging per word.

How to Free Text to Speech — Download as MP3 or WAV

  1. 1

    Download the voice model

    The first visit fetches about 86 MB and caches it in your browser. After that the tool works offline, and the download does not happen again.

  2. 2

    Type or paste your text

    Up to 5,000 characters per generation. Punctuation matters — full stops and commas are what the model uses to place pauses and shape intonation, so properly punctuated text sounds markedly better.

  3. 3

    Pick a voice and speed

    Fifteen voices across American and British English, listed best-first by the authors’ own quality grades, with thirteen rougher ones available behind a switch under Advanced. Press play on any of them to hear a sample before you commit. Speed runs from half to double, and changes pace without altering pitch.

  4. 4

    Generate and download

    Press generate — the button tells you roughly how long it will take — then play it back in the page and download it as an MP3 at your chosen bitrate or as a lossless WAV.

How it works

The voice comes from Kokoro, an 82-million-parameter text-to-speech model published under the Apache 2.0 licence. It is unusually small for the quality it reaches — around 86 MB, against the several gigabytes a typical cloud voice model needs — which is what makes it practical to send the model to you instead of sending your text to it. Kokoro does not read letters: text is first converted into phonemes, the individual sounds of speech, by a pronunciation dictionary of roughly 125,000 words with letter-to-sound rules for anything outside it. Those phonemes and a 256-number style vector describing the chosen voice go into the model, which returns a 24 kHz waveform. All of it happens in a Web Worker on your own machine, so a long script does not freeze the page and nothing you type is transmitted.

Supported formats

WAV, MP3, M4A, AAC, OGG and FLAC. Files are decoded by your browser, so anything it can play will work here. Output is exported as WAV.

Frequently asked questions

Is my text sent to a server?
No. The model is downloaded to your browser the first time you use the tool, and from then on the text you type is converted to speech on your own device. You can prove it: load the page once, turn off your Wi-Fi, and it still works. Nothing you write is transmitted, logged or stored.
Can I use the audio commercially?
Yes. Kokoro is released under the Apache 2.0 licence, which permits commercial use with no revenue ceiling and no separate registration, and the audio you generate is yours to use in videos, podcasts, adverts, games or client work. There is no watermark. This is worth checking with any text-to-speech tool — several popular open voice models are licensed for non-commercial use only, and several hosted services claim rights over the audio you generate on a free plan.
What languages does it support?
English only at the moment, in American and British accents. Kokoro’s main release is an English model; if you need another language, the transcription tool here handles over ninety, but speech generation is English for now.
How long does it take to generate?
Roughly twice as long as the speech itself, so about two minutes of waiting for a minute of audio, and around ten seconds for a short paragraph. The page estimates the wait on the button before you press it. That is the trade for running on your own device instead of a server farm: a cloud service returns audio faster, but only because the work is happening on someone else's hardware, with your text on it. A faster machine will beat these numbers and an older laptop or a phone will be slower.
How long can the text be?
Five thousand characters per generation, which is roughly five minutes of speech. Longer scripts are split at sentence boundaries automatically and joined back together, so you do not have to break them up yourself. There is also a daily allowance shared with the other free tools, measured in minutes of audio produced.
A word is pronounced wrong. Can I fix it?
Open “Show pronunciation” under the result and you will see the exact phonemes the model was given, which usually makes the cause obvious. The quickest workaround is to respell the word phonetically in your text — writing “Kokoro” as “Koh-koh-roh”, for example. Unusual names and technical jargon are the common cases, since they fall outside the pronunciation dictionary and get guessed by rule.
Why do the voices have letter grades?
They are the Kokoro authors’ own ratings, and they are worth paying attention to because quality varies a lot more between voices of one model than you might expect — within this one model the range runs from A to F+. The fifteen voices graded C or better are shown by default. The other thirteen are still there, behind a “show low-rated voices” switch under Advanced: they are noticeably rougher, with artefacts and odd stress, but you can hear every one of them before deciding, and one of them may be exactly the character you want.
Can it clone my own voice?
Not this tool — Kokoro generates from a fixed set of voices and has no voice-cloning capability at all. Voice cloning is a separate tool here, which records a few seconds of a voice and then speaks your text in it. It needs a browser with WebGPU and a much larger download, which is why the two are kept apart: this one runs anywhere and is the right choice whenever you just need a good voice rather than a particular person’s.
Does this use AI?
Yes. The voice is generated by Kokoro, a neural text-to-speech model, rather than stitched together from pre-recorded clips. The difference from a cloud service like ElevenLabs or Google’s Text-to-Speech is only where the model runs, which here is your own device.
Do I need an account?
No. The free tools work straight away with no sign-up. Creating a free account raises your daily processing limit from 20 to 60 minutes.
Is it free?
Yes. Every tool on this page is free to use, with a daily limit on total audio processed. There is no watermark and no trial period.

Other free audio tools