Free · no sign-up · runs on your device

Free text to speech with real AI voices.

Type or paste your text, pick a voice, and download a WAV. Fifteen natural voices, adjustable speed, no account, no watermark, no upload limit games. The AI model runs inside your own browser, so your words never touch a server.

0 / 2,000 characters

Voice

Speed

1.00x

Now imagine that in your voice.

That was one of our stock voices. Xport Studio takes a vocal you actually performed, with your own timing and phrasing, and converts it into any voice you have a model for. No length cap, no upload, fully offline on your machine.

Download Xport Studio, free Free to download · Pro is one time, never a subscription

The first run downloads the voice model once (about 86 MB) and your browser caches it. Everything after that is instant and offline.

These are our stock voices.

Xport Studio does the thing this page cannot: it takes a vocal you actually performed, with your timing, phrasing and emotion, and converts it into any voice you have a model for. Train your own, keep it on your machine, and nothing gets uploaded. Free to download.

Download Xport Studio

A text to speech tool that actually gives you the file

Most free text to speech sites make you sign up, cap you at a few minutes a day, stamp a watermark on the output, or simply refuse to let you download the audio at all. This one does none of that. You get a real WAV file, at full quality, for nothing, and you can generate as many as you want.

It is also genuinely private. The neural voice model is downloaded to your browser the first time you press generate, and from then on every word you type is converted on your own processor. There is no API call, no account, and no server that sees your script. If you disconnect from the internet after the first load, it keeps working.

Free text to speech with no sign-up

There is no account step here, and there is no email wall in front of the audio. Open the page, type, press generate. We can afford to leave it open because the work happens on your computer rather than ours, so a thousand people using it costs us the same as nobody using it. That is also why there is no daily minute counter: nothing is metered because nothing is being spent.

Download the audio as a WAV

Every generation gives you a real, uncompressed WAV file you can drag straight into a DAW, a video editor, or a podcast timeline. No watermark is mixed into it and no branded outro is appended. This is worth checking before you commit to any free text to speech site, because several of the popular ones will happily play audio at you and then refuse to hand over the file, sometimes even on their paid tiers. If you need a smaller file, run the WAV through our audio converter to get an MP3, M4A, FLAC, or OPUS.

Text to speech that runs offline

The first time you press generate, your browser downloads the voice model, roughly 86 MB, and keeps it cached. From that point on the tool is doing local inference: no request is sent, no server sees your script, and it keeps working if your connection drops. If you handle anything sensitive, unreleased lyrics, client scripts, internal documentation, medical or legal text, that distinction matters more than voice quality. Most text to speech services post your words to an API. This one cannot, because there is no API.

What you can use it for

Scratch vocals and reference hooks while you are writing, so you can hear how a line sits before you book studio time. Voiceover for a YouTube video, a TikTok, a Reel, or a course module before you hire talent. Narration for a demo or a walkthrough. Placeholder reads for a podcast intro. Audiobook-style playback of your own drafts, which is a fast way to catch clumsy sentences. Accessibility, if you would simply rather listen to something than read it. If you need it to sound like a specific person rather than a stock voice, that is what the desktop app is for.

The voices

Fifteen voices ship with the tool, split across American and British accents with male and female options in each: Iris, Nova, Lyra, Sky, Wren, and Sable on the American female side, Atlas, Onyx, Shadow, Echo, and Puck on the American male side, then Emma, Isla, George, and Fable for British English. Wren is the softest and reads well for calm narration, Onyx is the deepest, and Puck is the brightest. The speed slider runs from half speed to one and a half times, which is usually enough to match a video edit without pitching the voice.

More free audio tools

This is one of a dozen free in-browser tools we keep online, all of them private in the same way. The key and BPM finder reads the key and tempo of any track. The audio cutter and audio joiner trim and stitch files losslessly. The pitch and speed changer transposes without the chipmunk effect, and the recorder captures whatever your computer is playing. The full set is on the free tools page.

Questions

Is it really free, with no sign-up?

Yes. No account, no email, no trial timer, no watermark on the audio. The model runs on your device, which means it costs us nothing to serve, which is why we can leave it open.

Can I download the audio?

Yes, as a WAV. That is worth saying out loud because several popular text to speech services will not let you export audio even on a paid plan.

Is my text uploaded anywhere?

No. The only network request is the one-time download of the open voice model. After that your text is converted locally by your browser and never leaves your machine. We cannot see what you typed.

Why is the first generation slow?

The first press downloads the model, roughly 86 MB, and your browser caches it. After that, machines with WebGPU generate several times faster than real time. Older machines fall back to WebAssembly, which is slower but still works.

Can I make it sound like a specific artist or person?

Not here, deliberately. This page only offers our own stock voices, so no one's likeness is being cloned on our servers. Xport Studio, our desktop app, does voice conversion on a performance you recorded yourself, using a model you own, entirely offline.

What is the character limit?

2,000 characters per generation, which is roughly two minutes of speech, and you can run it as many times as you like. Longer scripts are better handled in the desktop app.

Which voices are included?

Fifteen, across American and British accents, male and female. They come from Kokoro, an open source 82M-parameter speech model, credited on our free tools page alongside the other open engines we use.

Can I use the audio commercially, on YouTube or in a podcast?

Yes. The voices come from Kokoro, which is released under the permissive Apache 2.0 licence, and we add no restriction of our own on top. There is no watermark and no attribution requirement for the audio you generate here.

Does it work on a phone?

It runs in a mobile browser, but the model download and the on-device processing are heavy, so a laptop or desktop is a much better experience. On a phone, expect the first load to take a while and generation to be slower.

Can I get an MP3 instead of a WAV?

Generate the WAV here, then drop it into our free audio converter, which will turn it into MP3, M4A, FLAC, OGG, or OPUS. That tool is also fully in-browser, so the file still never leaves your machine.

Is there a word or time limit?

2,000 characters per generation, about two minutes of speech, and there is no daily limit and no cap on how many times you run it. Longer pieces are simply done in sections.

Why does it need to download 86 MB?

That is the neural voice model itself. Cloud text to speech tools avoid the download by keeping the model on their servers and sending your text to them. We took the opposite trade: one bigger first load, then privacy and no limits forever after. Your browser caches it, so it only happens once.