Skip to content

What is Speech-to-Text (ASR)?

AI & Machine Learning, explained by the engineers who build it. Definition, how it works, use cases and common questions.

ASR definition

Speech-to-text, also called automatic speech recognition (ASR), is technology that converts spoken audio into written text. Modern systems use deep learning models, from open models such as OpenAI's Whisper and NVIDIA's Parakeet to cloud services from OpenAI, Google, AWS, Azure and Deepgram, to transcribe calls, meetings, voice notes and commands in real time or from recordings, across many languages and accents.

How speech recognition works

Audio is first converted into a representation the model can process, typically a spectrogram showing how frequencies change over time. An acoustic model, today usually a transformer or conformer network, maps these features to sounds and words, and a language model component predicts likely word sequences to resolve ambiguity, such as their versus there. End-to-end models like Whisper learn the whole mapping from audio to text directly from large amounts of transcribed audio.

Post-processing then adds punctuation, capitalization and number formatting, turning twenty five dollars into $25, and optionally speaker diarization, which labels who spoke when. Timestamps for each word let applications link text back to the exact moment in the recording.

Real-time vs batch transcription

Batch transcription processes complete recordings, such as recorded calls, podcasts or meetings, and can use larger models and more context for higher accuracy. Real-time or streaming transcription returns partial results within a fraction of a second as someone speaks, which is essential for live captions, voice assistants and voice AI agents, where any delay makes a conversation feel unnatural.

Streaming systems must balance speed and accuracy: early partial words may be revised as more audio arrives. They also need endpointing, detecting when a speaker has finished, which is surprisingly hard in noisy environments or with people who pause mid-sentence.

Accuracy, languages and accents

Accuracy is usually measured with word error rate (WER): the share of words substituted, inserted or deleted compared with a human transcript. It varies widely with audio quality, background noise, microphone distance, accents, crosstalk and vocabulary. Phone audio is harder than studio audio, and domain terms such as drug names or product codes are common failure points.

Improve results with custom vocabulary or phrase boosting, models tuned for your domain or language, better audio capture, and testing on recordings from your real users. Code-switching, such as Hindi mixed with English or Arabic with English, needs models and evaluation designed for it, since general benchmarks rarely reflect it.

Privacy shapes the architecture too. Recordings of calls and consultations are sensitive, so many teams prefer providers with regional hosting and no data retention, or self-host open models, and redact card numbers and other sensitive details from transcripts automatically.

Tools and business uses

Options range from open models you host yourself to managed APIs that add streaming, diarization and domain-specific models, usually priced per minute of audio. Common choices today, and the business uses they typically power across industries, include:

  • Open models: Whisper and its faster variants, plus newer models such as NVIDIA Parakeet and Canary, for self-hosted transcription
  • Cloud APIs: Google Cloud Speech-to-Text, Amazon Transcribe, Azure AI Speech, Deepgram and AssemblyAI
  • Contact centers: transcribing and analyzing calls for quality, compliance and sentiment
  • Meetings and media: notes, summaries, captions and searchable archives
  • Healthcare: clinical documentation from doctor and patient conversations
  • Voice interfaces: commands, dictation and voice agents

ASR: common questions

Something else on your mind? Ask a consultant and get a reply within one business day.

What is the difference between speech-to-text and voice recognition?

Speech-to-text, or speech recognition, converts what was said into text. Voice recognition, or speaker recognition, identifies who is speaking based on voice characteristics, for example for authentication. Some systems do both, such as transcribing a meeting while labeling each speaker, but they are different technologies.

Is Whisper good enough for production?

For many uses, yes. Whisper is accurate across many languages and robust to noise, and optimized versions run fast on GPUs. It does not stream natively, may invent text during long silences and lacks built-in diarization, so production systems often add voice activity detection, streaming wrappers and speaker labeling.

Is speech data personal data?

Usually, yes. Recordings and transcripts often contain names, account details and health or financial information, and voiceprints can count as biometric data under some laws. Obtain consent where required, minimize retention, redact sensitive fields from transcripts and choose providers whose data handling and hosting regions meet your obligations.

Keep exploring the ai & machine learning glossary

Need ASR in your product?

A solutions consultant replies within one business day with next steps, a rough estimate and a suggested team.