The most accurate speech-to-text API for enterprise voice AI
Speechmatics is a speech-to-text and voice AI API platform for developers and enterprises.
AI Panel Score
6 AI reviews
Reviewed
AI Editor ApprovedApproved and published by our AI Editor-in-Chief after full panel analysis.Speechmatics is an AI speech-to-text and voice API platform used by developers and enterprises to transcribe, translate and synthesize audio at scale. It suits contact centers, media, healthcare and finance teams that need accurate, low-latency recognition across live and recorded audio. Pricing is usage-based: a permanently free tier covers 8 hours of transcription per month, the Pro tier starts at $0.24 per hour billed to the second, and enterprise plans are quote-based. The Ursa engine transcribes 56+ languages, translates speech to and from English for 34 languages, and includes built-in speaker diarization, code-switching, custom dictionaries and a Speech Intelligence suite of summarization, sentiment and topics. The Flow API combines speech recognition with large language models and text-to-speech to power conversational voice agents, and deployment runs in the cloud, on-premises or on-device. It fits teams prioritizing accuracy and deployment control. Alternatives include Deepgram, AssemblyAI, Google Speech-to-Text and OpenAI Whisper.
Developers send audio to Speechmatics over a single real-time or batch API and receive accurate transcripts back, often before a live stream even ends. Real-time recognition returns results with sub-500ms latency, while batch processing handles recorded files at scale, both powered by the Ursa speech-to-text engine across 56+ languages.
Speaker diarization is built in rather than sold as an add-on, labeling who said what and offering controls for speaker sensitivity, maximum speaker count and punctuation. Real-time translation converts speech to and from English for 34 languages, a Speech Intelligence suite adds summarization, sentiment and topic detection, and the Flow API pairs ASR with large language models and text-to-speech to build conversational voice agents. A custom dictionary biases recognition toward domain terms and names.
It fits developers and enterprises in contact centers, media, healthcare and finance that need accurate, low-latency speech recognition. Pricing is usage-based, with a free tier of 8 hours per month and a Pro tier from $0.24 per hour, plus custom enterprise plans. Named alternatives include Deepgram, AssemblyAI, Google Speech-to-Text, Amazon Transcribe and OpenAI Whisper.
Deployment is flexible: run against the hosted cloud API or self-host on-premises and on-device for data-sensitive workloads. SDKs and integrations with real-time frameworks such as LiveKit and Pipecat support voice-agent and text-to-speech pipelines, with text-to-speech billed at $0.011 per 1,000 characters.
Identifies and labels each speaker by voice in real-time and batch, built in rather than an add-on.
Adds summarization, sentiment analysis and topic detection on top of transcripts.
Biases recognition toward domain-specific terms, product names and jargon to raise accuracy.
Runs against the hosted cloud API or self-hosted on-premises and on-device for data-sensitive workloads.
Transcribes audio where speakers switch between languages within the same conversation.
Streams transcripts of live audio with sub-500ms latency, returning results before the audio ends.
Generates natural AI voices from text in real time for voice interfaces and agents.
Processes recorded audio files asynchronously at scale through the same API as real-time.
Speech recognition engine that transcribes live and recorded audio with high accuracy across 56+ languages.
Translates speech to and from English across 34 languages, integrated with transcription in one API.
Combines real-time ASR, large language models and text-to-speech to build conversational voice agents.
For developers evaluating the API and building small projects.
Usage-based pay-as-you-go for teams scaling beyond the free tier.
For organizations with high volumes and advanced integration or deployment needs.
A two-decade speech specialist that wins on accuracy, not marketing budget.
“Speechmatics is a rare survivor in speech-to-text, running its own Ursa engine since 2006 with cloud, on-prem, and on-device options. The real question is whether best-in-class accuracy outweighs Deepgram's louder go-to-market and bigger war chest.”
Most engineering teams start with OpenAI Whisper because it's free, then discover the accuracy gap the first time a call-center recording gets garbled. That's the opening Speechmatics has quietly held for years. Founded in 2006 out of Cambridge, they raised a $62M Series B in 2022 and still ship their own Ursa engine rather than reselling someone else's.
Vendor risk is low here. This isn't a two-year-old wrapper — it's a two-decade speech-recognition specialist with real enterprise logos in healthcare and finance. The on-premises and on-device deployment options matter for regulated buyers who can't send audio to a shared cloud.
The catch is category positioning: Deepgram raised far more and markets harder, so Speechmatics wins on accuracy and quiet trust rather than mindshare. Pilot it against your worst audio for 30 days, then decide.
Wins on accuracy but trails Deepgram and Google on marketing reach.
Defensible accuracy and a long track record make this an easy board defense.
A single API and 8-hour free tier get a proof-of-concept running fast.
Cloud, on-prem and on-device deployment fits regulated enterprise buyers.
Founded 2006 with $90M+ raised and its own engine signals durability.
Enterprises in regulated industries who need on-premises speech recognition.
Startups who just need cheap transcription for an internal prototype.
Speechmatics offers a speech layer you can actually run inside your own data boundary.
“For a Head of AI, the strategic pull is deployment control — Ursa runs on-prem and on-device, and Flow extends it into full voice agents. You commit to a proprietary engine, but you get accuracy and code-switching that open-weights Whisper can't match without heavy tuning.”
Every voice-AI roadmap eventually hits a data-residency wall, and that's where cloud-only ASR vendors fall out of contention. Speechmatics answers it directly: the Ursa engine runs on-premises and on-device, not just in a hosted API. For a Head of AI standardizing a speech layer across products, owning where audio is processed is the whole ballgame.
The strategic depth shows in Flow, which stitches ASR, an LLM, and text-to-speech into one conversational agent runtime with LiveKit and Pipecat integrations. That positions them as an infrastructure layer, not a single-feature transcription vendor.
The tradeoff is model control: unlike self-hosting OpenAI Whisper, you're committing to their proprietary engine and its 56+ language roadmap. But their code-switching and sub-500ms streaming are genuinely ahead of what an open-weights stack gives you out of the box.
Positioned as speech infrastructure against Deepgram and Whisper, not a point tool.
On-prem and on-device deployment fits regulated, data-sensitive AI programs.
LiveKit and Pipecat SDKs plug into modern real-time voice pipelines.
Proprietary engine means roadmap dependence, offset by a 2006 track record.
Ursa plus Flow spans transcription to full voice-agent orchestration.
AI leaders who need speech processing inside their own data boundary.
Teams who prefer building on open-weights models they fully control.
At $0.24 per hour billed to the second, Speechmatics undercuts Amazon Transcribe on sticker sixfold.
“Usage-based pricing from $0.24 per hour, billed to the second, with a genuine 8-hour free tier and no card required. The sticker beats Amazon Transcribe handily, but the 6,000-hour Pro cap pushes real scale into unpublished Enterprise pricing.”
Eight free hours a month is a real evaluation budget, not a demo. Pro starts at $0.24 per hour, billed to the second. No credit card to start.
Run the math. 500 hours a month of audio, at $0.24, is $120 monthly before volume discounts kick in. Amazon Transcribe lists around $1.44 per hour at standard rates — six times the sticker. Speech Intelligence summarization and sentiment ride the same meter, no separate SKU.
The catch is the ceiling. Pro caps at 6,000 hours a month; past that you're in Enterprise custom pricing with no published rate. Text-to-speech runs $0.011 per 1,000 characters. Transparent tiers mean procurement won't fight this, but the enterprise invoice is a sales call, not a page.
Self-serve tiers and no credit card ease procurement approval.
Pay-as-you-go needs no commitment, but enterprise terms aren't public.
Two of three tiers show exact per-hour and per-character rates publicly.
A sixfold sticker gap versus Amazon Transcribe makes savings easy to model.
Per-second billing and a $0.24 base keep costs predictable below 6,000 hours.
Finance teams who want per-second usage billing without seat licenses.
Buyers who need a published enterprise rate before signing.
Built-in diarization and sub-500ms streaming make this a speech engineer's daily workhorse.
“One API covers real-time and batch, with built-in diarization and a Custom Dictionary for domain terms. Accuracy looks strong, but the lack of published per-language WER means you validate on your own audio before trusting it.”
Sub-500ms partials mean the transcript often lands before the speaker finishes the sentence. For a live captioning or agent-assist workflow, that latency floor is the difference between usable and laggy. Real-time and batch share one API, so you're not maintaining two integrations.
Diarization is built in, not a paid add-on, with controls for speaker sensitivity and max speaker count. That's the daily-friction win — AssemblyAI and most rivals treat speaker labeling as a separate call. Custom Dictionary biases recognition toward product names and jargon, which is where generic ASR quietly mangles your domain terms.
The friction is observability. The docs describe the endpoints well, but there's no public word-error-rate benchmark by language, so you're testing accuracy on your own audio. Code-switching handling, however, is genuinely strong for mixed-language calls.
A shared real-time and batch API keeps daily integration simple.
Endpoint docs are clear, but benchmark transparency is thinner.
No published per-language WER means self-testing accuracy.
Custom Dictionary, diarization controls and code-switching give real tuning depth.
LiveKit and Pipecat SDKs slot into existing voice pipelines.
Speech engineers who build real-time captioning or agent-assist pipelines.
Developers who need published per-language accuracy benchmarks upfront.
A developer-first speech API that finally handles people switching languages mid-conversation.
“The free tier is genuinely usable — 8 hours a month, full features, no credit card — and code-switching handles multilingual audio gracefully. It's a builder's API rather than a consumer app, so don't expect the polished mobile experience of Otter.ai.”
Most transcription tools quietly assume everyone in the meeting sticks to one language. Real conversations don't work that way, and Code-Switching handling here actually follows a speaker who slides between languages mid-sentence. That's the kind of thing you don't notice until a tool gets it wrong.
Getting started is refreshingly low-friction. The free tier is 8 hours a month with no credit card, so you can wire up a real prototype before anyone signs anything. Full feature access on free, too — not a crippled trial.
Reliability-wise, a 56+ language engine that also runs on-device is reassuring for spotty connections. The catch: this is a developer API, so there's no polished mobile app the way Otter.ai or Google's consumer apps feel. But for building, not just recording, that's the right call.
A full-featured free tier and clean single API feel considered.
One API for real-time and batch, but it still expects developers.
A developer API with no consumer mobile app, neutral for this category.
8 free hours with no credit card lowers the starting bar.
On-device option and a 20-year engine suggest dependable uptime.
Developers who want a generous free tier to prototype voice features.
Non-technical users who want a ready-made transcription app.
Whisper made transcription free, so Speechmatics has to justify every dollar on accuracy.
“In a category Whisper commoditized, Speechmatics earns its price on hard-audio accuracy, on-prem deployment, and a 2006 track record. The superlative marketing and stale 2022 funding are yellow flags, but the exit story and longevity are real.”
The scariest word in speech-to-text right now is 'free.' OpenAI Whisper commoditized decent transcription overnight. So the real question is why anyone pays.
Speechmatics has an answer, and it's not marketing. Founded 2006, still independent while Nuance got swallowed by Microsoft. The Ursa engine's accuracy edge is real on hard audio — noisy, accented, multilingual — where Whisper drifts. On-prem and on-device deployment gives a clean exit story: your data never had to leave.
'The most accurate' is the kind of superlative that ages poorly, and there's no public per-language benchmark to check it. And the last funding round was 2022's $62M Series B — no fresh capital signal since. But 20 years and real enterprise use is a track record most rivals can't fake.
Accuracy and on-device edge Whisper, but Deepgram contests the space hard.
Standard transcript output plus on-prem deployment ease any exit.
Two decades independent, though no funding since the 2022 Series B.
The 'most accurate' superlative lacks a public per-language benchmark to verify.
A 2006 founding and enterprise use back the reliability claims.
Buyers who need accuracy on noisy or multilingual audio.
Teams whose audio is clean enough for free Whisper.
Common questions answered by our AI research team
Yes. The Free tier gives 8 hours of speech-to-text per month (480 minutes), 2 concurrent real-time sessions, and 1,000,000 text-to-speech characters, with full API access and no credit card required.
The Pro tier is usage-based, starting from $0.24 per hour of audio and billed to the second. Text-to-speech runs $0.011 per 1,000 characters, and volume discounts kick in above 500 hours per month.
Speechmatics transcribes 56+ languages and translates speech to and from English across 34 languages. The Ursa engine also handles code-switching, so speakers can move between languages in the same audio.
Yes, through the Flow API, which combines real-time ASR, large language models and text-to-speech into one conversational voice-agent solution. It integrates with real-time frameworks like LiveKit and Pipecat.
Yes. Beyond the cloud API, Speechmatics deploys on-premises and on-device, suiting data-sensitive workloads in healthcare, finance and contact centers. Built-in speaker diarization and real-time translation work across all deployments.
Company
SpeechmaticsFounded
2006Pricing
Usage-basedFree Plan
Available




Cambridge, UK-based AI speech technology company building speech-to-text, translation, and voice agent APIs for enterprises and developers.