What is speech to text (STT)?
Speech to text (STT), also called automatic speech recognition (ASR), is a technology that converts spoken language into written text. It powers transcription, automatic subtitles, dictation and voice assistants. Accuracy depends mostly on audio quality, accents and vocabulary. On wawie, speech to text drives automatic subtitles you can edit, translate and export for YouTube, TikTok or Instagram.
Create a free account 5,000 free credits. No card needed.
Updated September 27, 2026
How does speech to text work?
Speech to text first turns the audio into a compact picture of its frequencies over time, usually a spectrogram. A neural network, trained on large amounts of transcribed speech, maps those features to likely sounds, word pieces or characters. A language model, built into the network or added after it, uses context to pick the most probable words, so the system can tell to, two and too apart. Post-processing then adds punctuation, capital letters and number formatting, and attaches timestamps to words or phrases. For subtitles, the text is split into short caption lines timed to the speech and exported in formats such as SRT or VTT.
What is speech to text used for?
Speech to text is used whenever spoken words need to become searchable, readable text. Video creators use it to generate subtitles and captions, which help viewers who watch without sound and make content accessible to deaf and hard-of-hearing people. Teams transcribe meetings, interviews and lectures so they can search and quote them. Writers and professionals use dictation to type by voice. Speech to text is also the first step of AI dubbing, where the transcript is translated before being voiced again. A transcript also gives search engines and AI assistants text they can index.
- Subtitles and captions for YouTube, TikTok and Instagram
- Meeting, interview and lecture transcripts
- Dictation and voice typing
- The first step of AI dubbing
What affects speech to text accuracy?
Speech to text accuracy is usually measured with word error rate (WER), the share of words that are wrong, missing or added compared with a correct transcript. Clean audio from one speaker close to the microphone gives the best results. Accuracy drops with background noise, music, echo, people talking over each other, accents the model has rarely heard and specialised vocabulary such as brand names, medical terms or jargon. Fast speech and low-quality recordings add errors too. For anything you publish, review the transcript and fix names and key terms before exporting captions.
Examples of speech to text
- A YouTuber generates captions for a tutorial, corrects two product names and exports an SRT file.
- A TikTok creator burns styled captions into a short video for viewers who watch on mute.
- A researcher transcribes recorded interviews to search and quote them in a report.
- A phone's dictation feature types a message as you speak.
Frequently asked questions
What is the difference between speech to text and text to speech?
They work in opposite directions. Speech to text (STT) listens to spoken audio and writes it down, producing a transcript or subtitles. Text to speech (TTS) takes written text and reads it aloud in a synthetic voice. Speech to text is about understanding speech, while text to speech is about producing it. The two are often combined, for example in AI dubbing, where a video is transcribed, translated and then voiced again in another language.
How accurate is speech to text?
On clear speech in widely supported languages, modern speech to text is highly accurate, and most remaining errors involve names, jargon, numbers and overlapping speakers. Accuracy falls with background noise, music, echo, fast speech and unfamiliar accents. It is measured with word error rate (WER): the lower the rate, the better. Captions for published videos should always be reviewed. On wawie, every caption is editable before export, and you can test it free with 5,000 credits, no card needed.
What are SRT and VTT subtitle files?
SRT (SubRip) and VTT (WebVTT) are two of the most common subtitle file formats. Both are plain text files containing a list of caption cues, each with a start time, an end time and the text to display. SRT is supported by almost every video platform and editor. WebVTT is the web standard for HTML5 video and supports extra styling and positioning. You upload either file alongside your video, so viewers can turn captions on or off.
What is the difference between captions and subtitles?
In the United States, captions usually mean same-language text for viewers who cannot hear the audio, so they also describe relevant sounds such as music or laughter and may identify speakers. Subtitles usually mean a transcript or translation of the dialogue only, for viewers who can hear. In the UK, the word subtitles is often used for both. Closed captions can be switched off, while open or burned-in captions are part of the video image and always visible.
Related terms
- Text to speech (TTS): Text to speech converts written text into natural-sounding spoken audio, used for voiceovers, narration, accessibility and voice assistants.
- AI dubbing: AI dubbing translates the speech in a video and re-voices it in another language, timed to the original and optionally in the original voice.
- Voice isolation: Voice isolation separates a voice from background noise, music and room echo, leaving clean speech you can publish, transcribe or dub.
Create a free account 5,000 free credits. No card needed.