What is text to speech (TTS)?
Text to speech (TTS), also called speech synthesis, is a technology that converts written text into spoken audio. Modern TTS uses neural networks trained on recorded speech, so voices follow natural rhythm and intonation. It powers screen readers, voice assistants, voiceovers and narration. On wawie, you can turn a script into speech with 1000+ voices in about 30 languages.
Create a free account 5,000 free credits. No card needed.
Updated September 27, 2026
How does text to speech work?
Text to speech works in stages. First, text normalisation rewrites numbers, dates and abbreviations as words, so 10 km becomes ten kilometres. Next, the system converts words into phonemes, the sound units of a language, and predicts prosody: where to pause, which syllables to stress and how the pitch should rise and fall. A neural acoustic model then generates a representation of the sound, often a spectrogram, and a vocoder turns it into an audio waveform. Newer end-to-end models perform several of these steps in one network. Older systems stitched together recorded fragments of speech, which could sound uneven where the pieces joined.
What is text to speech used for?
Text to speech is used wherever written content needs a voice without a recording session. It gives blind and low-vision users access to screens and helps people who find reading difficult to listen instead. Creators use it for video voiceovers, course narration and podcast segments. Businesses use it for phone menus, product demos and announcements that change often, because a line can be edited and regenerated in seconds instead of re-recorded. Publishers use it to offer audio versions of articles and books. With a large voice library, you can also test several voices on the same script before choosing one.
- Accessibility and screen readers
- Voiceovers for video and social media
- E-learning and training narration
- Audiobooks and audio articles
- Phone menus and announcements
What affects the quality of text to speech?
The quality of text to speech depends on the voice, the language and the text itself. Some voices sound natural in one language and flat in another, so preview before you commit. Names, brands, acronyms and rare words cause the most common errors; a pronunciation dictionary or phonetic spelling usually fixes them. Punctuation controls pacing: commas and short sentences create natural pauses. Very long passages can drift in tone, so generating paragraph by paragraph gives you more control. Finally, check the licence: commercial rights to generated audio depend on each provider's terms, and cloning a real person's voice requires their consent.
Examples of text to speech
- A screen reader reads a web page aloud so a blind user can navigate it.
- A YouTube creator pastes a script, picks a voice and exports the narration as an MP3 for the edit.
- An e-learning team changes one sentence in a course and regenerates only that line instead of booking a new recording session.
- A navigation app announces each turn in the driver's chosen language.
Frequently asked questions
What is the difference between text to speech and speech to text?
Text to speech and speech to text are opposite processes. Text to speech (TTS) takes written text and produces spoken audio, for example a voiceover from a script. Speech to text (STT), also called speech recognition, takes spoken audio and produces written text, for example a transcript or subtitles from a video. Many workflows use both: AI dubbing transcribes speech with STT, translates the text, then voices the translation with TTS.
Can text to speech sound like a real person?
Yes. Modern neural text to speech can sound very close to a human speaker, with natural pauses, breathing and intonation, especially in short and medium-length passages. Listeners are more likely to notice a synthetic voice in long, emotional or highly expressive content, where a human performer still adds nuance. Quality also varies by voice and by language, so it is worth previewing several voices. Reproducing a specific real person's voice requires voice cloning and that person's explicit consent.
Can I use text to speech audio commercially?
It depends on the service. Licence terms vary from one provider to another, so read them before you publish. On wawie, content you generate from your own text is yours to use, within the terms of the underlying model providers, such as ElevenLabs, MiniMax and Fish Audio. You remain responsible for the rights to the text you supply, and you need explicit consent before using anyone else's voice through cloning. Commercial use is included from the Wawie Start plan at EUR 4.99 a month.
How can I try text to speech for free?
Many text to speech services offer a free tier. On wawie, you can create a free account with 5,000 one-time credits and no card needed. Paste a script, choose among 1000+ voices in about 30 languages, preview them and download the result as MP3 or WAV. Credits are spent on actual usage, per character read or per second of audio produced. Paid plans start at EUR 4.99 a month when you need more.
Related terms
- Voiceover: A voiceover is narration recorded separately and played over video, film, ads or presentations, by a speaker who is usually off screen.
- Voice cloning: Voice cloning builds a synthetic copy of a specific person's voice from recordings, so new speech can be generated in that voice.
- Speech to text (STT): Speech to text converts spoken audio into written text, powering transcripts, automatic subtitles, dictation and voice assistants.
Create a free account 5,000 free credits. No card needed.