Speechmatics logo

Speechmatics Review

Visit

The most accurate speech-to-text API for enterprise voice AI

Speechmatics is a speech-to-text and voice AI API platform for developers and enterprises.

AI Panel Score

8.0/10

6 AI reviews

Reviewed

AI Editor Approved

What is Speechmatics?

Speechmatics is an AI speech-to-text and voice API platform used by developers and enterprises to transcribe, translate and synthesize audio at scale. It suits contact centers, media, healthcare and finance teams that need accurate, low-latency recognition across live and recorded audio. Pricing is usage-based: a permanently free tier covers 8 hours of transcription per month, the Pro tier starts at $0.24 per hour billed to the second, and enterprise plans are quote-based. The Ursa engine transcribes 56+ languages, translates speech to and from English for 34 languages, and includes built-in speaker diarization, code-switching, custom dictionaries and a Speech Intelligence suite of summarization, sentiment and topics. The Flow API combines speech recognition with large language models and text-to-speech to power conversational voice agents, and deployment runs in the cloud, on-premises or on-device. It fits teams prioritizing accuracy and deployment control. Alternatives include Deepgram, AssemblyAI, Google Speech-to-Text and OpenAI Whisper.

About Speechmatics

Developers send audio to Speechmatics over a single real-time or batch API and receive accurate transcripts back, often before a live stream even ends. Real-time recognition returns results with sub-500ms latency, while batch processing handles recorded files at scale, both powered by the Ursa speech-to-text engine across 56+ languages.

Speaker diarization is built in rather than sold as an add-on, labeling who said what and offering controls for speaker sensitivity, maximum speaker count and punctuation. Real-time translation converts speech to and from English for 34 languages, a Speech Intelligence suite adds summarization, sentiment and topic detection, and the Flow API pairs ASR with large language models and text-to-speech to build conversational voice agents. A custom dictionary biases recognition toward domain terms and names.

It fits developers and enterprises in contact centers, media, healthcare and finance that need accurate, low-latency speech recognition. Pricing is usage-based, with a free tier of 8 hours per month and a Pro tier from $0.24 per hour, plus custom enterprise plans. Named alternatives include Deepgram, AssemblyAI, Google Speech-to-Text, Amazon Transcribe and OpenAI Whisper.

Deployment is flexible: run against the hosted cloud API or self-host on-premises and on-device for data-sensitive workloads. SDKs and integrations with real-time frameworks such as LiveKit and Pipecat support voice-agent and text-to-speech pipelines, with text-to-speech billed at $0.011 per 1,000 characters.

Features

Analytics

  • Speaker Diarization

    Identifies and labels each speaker by voice in real-time and batch, built in rather than an add-on.

  • Speech Intelligence

    Adds summarization, sentiment analysis and topic detection on top of transcripts.

Customization

  • Custom Dictionary

    Biases recognition toward domain-specific terms, product names and jargon to raise accuracy.

Deployment

  • Flexible Deployment

    Runs against the hosted cloud API or self-hosted on-premises and on-device for data-sensitive workloads.

Language

  • Code-Switching

    Transcribes audio where speakers switch between languages within the same conversation.

Real-Time

  • Real-Time Streaming

    Streams transcripts of live audio with sub-500ms latency, returning results before the audio ends.

Speech Synthesis

  • Text-to-Speech

    Generates natural AI voices from text in real time for voice interfaces and agents.

Transcription

  • Batch Transcription

    Processes recorded audio files asynchronously at scale through the same API as real-time.

  • Ursa Speech-to-Text

    Speech recognition engine that transcribes live and recorded audio with high accuracy across 56+ languages.

Translation

  • Real-Time Translation

    Translates speech to and from English across 34 languages, integrated with transcription in one API.

Voice Agents

  • Flow Voice Agents API

    Combines real-time ASR, large language models and text-to-speech to build conversational voice agents.

Preview

Speechmatics desktop previewSpeechmatics mobile preview

Pricing Plans

Free

Free

For developers evaluating the API and building small projects.

  • 8 hours of speech-to-text per month (480 minutes)
  • 2 concurrent real-time sessions
  • 1,000,000 text-to-speech characters per month (~20 hours)
  • Full access to the feature suite
  • API access, no credit card required

Pro

$0/usage

Usage-based pay-as-you-go for teams scaling beyond the free tier.

  • Speech-to-text from $0.24 per hour, billed to the second
  • Text-to-speech at $0.011 per 1,000 characters
  • 50 concurrent real-time sessions
  • 10 batch file jobs per second
  • Volume discounts above 500 hours per month
  • Email support, capped at 6,000 hours per month

Enterprise

Contact sales

For organizations with high volumes and advanced integration or deployment needs.

  • Custom volume pricing
  • On-premises and on-device deployment
  • Advanced security and compliance
  • Dedicated support and SLAs
  • Bespoke integration support

AI Panel Reviews

The Decision Maker

The Decision Maker

Strategic bet, vendor viability, timing, adoption approval
8.2/10

A two-decade speech specialist that wins on accuracy, not marketing budget.

Speechmatics is a rare survivor in speech-to-text, running its own Ursa engine since 2006 with cloud, on-prem, and on-device options. The real question is whether best-in-class accuracy outweighs Deepgram's louder go-to-market and bigger war chest.

Most engineering teams start with OpenAI Whisper because it's free, then discover the accuracy gap the first time a call-center recording gets garbled. That's the opening Speechmatics has quietly held for years. Founded in 2006 out of Cambridge, they raised a $62M Series B in 2022 and still ship their own Ursa engine rather than reselling someone else's.

Vendor risk is low here. This isn't a two-year-old wrapper — it's a two-decade speech-recognition specialist with real enterprise logos in healthcare and finance. The on-premises and on-device deployment options matter for regulated buyers who can't send audio to a shared cloud.

The catch is category positioning: Deepgram raised far more and markets harder, so Speechmatics wins on accuracy and quiet trust rather than mindshare. Pilot it against your worst audio for 30 days, then decide.

Competitive Positioning7.7

Wins on accuracy but trails Deepgram and Google on marketing reach.

Reputation Risk8.2

Defensible accuracy and a long track record make this an easy board defense.

Speed to Value7.8

A single API and 8-hour free tier get a proof-of-concept running fast.

Strategic Fit8.0

Cloud, on-prem and on-device deployment fits regulated enterprise buyers.

Vendor Viability8.5

Founded 2006 with $90M+ raised and its own engine signals durability.

Pros

  • Two decades of speech-recognition focus with its own Ursa engine, not a reseller.
  • On-premises and on-device deployment suits healthcare and finance data rules.
  • Transparent usage pricing from $0.24 per hour with a real free tier.
  • 56+ language coverage with built-in diarization and translation.

Cons

  • Smaller marketing footprint than Deepgram or Google Speech-to-Text.
  • Last major raise was the 2022 Series B, with no newer round public.

Right for

Enterprises in regulated industries who need on-premises speech recognition.

Avoid if

Startups who just need cheap transcription for an internal prototype.

The Domain Strategist

The Domain Strategist

Craft and strategy in the product's domain — adapts identity per category, same lens
8.2/10

Speechmatics offers a speech layer you can actually run inside your own data boundary.

For a Head of AI, the strategic pull is deployment control — Ursa runs on-prem and on-device, and Flow extends it into full voice agents. You commit to a proprietary engine, but you get accuracy and code-switching that open-weights Whisper can't match without heavy tuning.

Every voice-AI roadmap eventually hits a data-residency wall, and that's where cloud-only ASR vendors fall out of contention. Speechmatics answers it directly: the Ursa engine runs on-premises and on-device, not just in a hosted API. For a Head of AI standardizing a speech layer across products, owning where audio is processed is the whole ballgame.

The strategic depth shows in Flow, which stitches ASR, an LLM, and text-to-speech into one conversational agent runtime with LiveKit and Pipecat integrations. That positions them as an infrastructure layer, not a single-feature transcription vendor.

The tradeoff is model control: unlike self-hosting OpenAI Whisper, you're committing to their proprietary engine and its 56+ language roadmap. But their code-switching and sub-500ms streaming are genuinely ahead of what an open-weights stack gives you out of the box.

Category Positioning8.0

Positioned as speech infrastructure against Deepgram and Whisper, not a point tool.

Domain Fit8.4

On-prem and on-device deployment fits regulated, data-sensitive AI programs.

Integration Surface8.1

LiveKit and Pipecat SDKs plug into modern real-time voice pipelines.

Long-term Implications7.9

Proprietary engine means roadmap dependence, offset by a 2006 track record.

Strategic Depth8.3

Ursa plus Flow spans transcription to full voice-agent orchestration.

Pros

  • Ursa runs on-premises and on-device, rare among accurate ASR vendors.
  • Flow unifies ASR, LLM, and text-to-speech into one agent runtime.
  • Strong code-switching and 34-language translation in a single API.
  • 56+ languages give broad global coverage for multi-market products.

Cons

  • Proprietary engine creates roadmap dependence versus open-weights alternatives.
  • No public funding round since 2022 to signal aggressive R&D scaling.

Right for

AI leaders who need speech processing inside their own data boundary.

Avoid if

Teams who prefer building on open-weights models they fully control.

The Finance Lead

The Finance Lead

Money, total cost of ownership, contracts, procurement math
8.1/10

At $0.24 per hour billed to the second, Speechmatics undercuts Amazon Transcribe on sticker sixfold.

Usage-based pricing from $0.24 per hour, billed to the second, with a genuine 8-hour free tier and no card required. The sticker beats Amazon Transcribe handily, but the 6,000-hour Pro cap pushes real scale into unpublished Enterprise pricing.

Eight free hours a month is a real evaluation budget, not a demo. Pro starts at $0.24 per hour, billed to the second. No credit card to start.

Run the math. 500 hours a month of audio, at $0.24, is $120 monthly before volume discounts kick in. Amazon Transcribe lists around $1.44 per hour at standard rates — six times the sticker. Speech Intelligence summarization and sentiment ride the same meter, no separate SKU.

The catch is the ceiling. Pro caps at 6,000 hours a month; past that you're in Enterprise custom pricing with no published rate. Text-to-speech runs $0.011 per 1,000 characters. Transparent tiers mean procurement won't fight this, but the enterprise invoice is a sales call, not a page.

Billing & Procurement8.1

Self-serve tiers and no credit card ease procurement approval.

Contract Flexibility7.6

Pay-as-you-go needs no commitment, but enterprise terms aren't public.

Pricing Transparency8.5

Two of three tiers show exact per-hour and per-character rates publicly.

ROI Clarity8.0

A sixfold sticker gap versus Amazon Transcribe makes savings easy to model.

Total Cost of Ownership8.0

Per-second billing and a $0.24 base keep costs predictable below 6,000 hours.

Pros

  • Free tier gives 8 hours monthly with full feature access.
  • Per-second billing avoids rounding waste on short clips.
  • $0.24 per hour undercuts hyperscaler ASR sticker pricing sharply.
  • Two tiers fully priced without a sales call.

Cons

  • Enterprise pricing is custom with no published rate.
  • Pro support caps at 6,000 hours per month.

Right for

Finance teams who want per-second usage billing without seat licenses.

Avoid if

Buyers who need a published enterprise rate before signing.

The Domain Practitioner

The Domain Practitioner

Daily hands-on reality in the product's domain — adapts identity per category, same lens
8.0/10

Built-in diarization and sub-500ms streaming make this a speech engineer's daily workhorse.

One API covers real-time and batch, with built-in diarization and a Custom Dictionary for domain terms. Accuracy looks strong, but the lack of published per-language WER means you validate on your own audio before trusting it.

Sub-500ms partials mean the transcript often lands before the speaker finishes the sentence. For a live captioning or agent-assist workflow, that latency floor is the difference between usable and laggy. Real-time and batch share one API, so you're not maintaining two integrations.

Diarization is built in, not a paid add-on, with controls for speaker sensitivity and max speaker count. That's the daily-friction win — AssemblyAI and most rivals treat speaker labeling as a separate call. Custom Dictionary biases recognition toward product names and jargon, which is where generic ASR quietly mangles your domain terms.

The friction is observability. The docs describe the endpoints well, but there's no public word-error-rate benchmark by language, so you're testing accuracy on your own audio. Code-switching handling, however, is genuinely strong for mixed-language calls.

Day-3 Reality8.0

A shared real-time and batch API keeps daily integration simple.

Documentation Practitioner-Fit7.8

Endpoint docs are clear, but benchmark transparency is thinner.

Friction Surface7.6

No published per-language WER means self-testing accuracy.

Power-User Depth8.2

Custom Dictionary, diarization controls and code-switching give real tuning depth.

Workflow Integration8.2

LiveKit and Pipecat SDKs slot into existing voice pipelines.

Pros

  • Sub-500ms real-time partials support live captioning workflows.
  • One API serves both real-time streaming and batch jobs.
  • Built-in diarization with speaker sensitivity and count controls.
  • Custom Dictionary raises accuracy on domain-specific vocabulary.

Cons

  • No published per-language word-error-rate benchmarks.
  • Accuracy validation falls on your own test audio.

Right for

Speech engineers who build real-time captioning or agent-assist pipelines.

Avoid if

Developers who need published per-language accuracy benchmarks upfront.

The Power User

The Power User

Daily human experience, onboarding, polish, learning curve, reliability
8.0/10

A developer-first speech API that finally handles people switching languages mid-conversation.

The free tier is genuinely usable — 8 hours a month, full features, no credit card — and code-switching handles multilingual audio gracefully. It's a builder's API rather than a consumer app, so don't expect the polished mobile experience of Otter.ai.

Most transcription tools quietly assume everyone in the meeting sticks to one language. Real conversations don't work that way, and Code-Switching handling here actually follows a speaker who slides between languages mid-sentence. That's the kind of thing you don't notice until a tool gets it wrong.

Getting started is refreshingly low-friction. The free tier is 8 hours a month with no credit card, so you can wire up a real prototype before anyone signs anything. Full feature access on free, too — not a crippled trial.

Reliability-wise, a 56+ language engine that also runs on-device is reassuring for spotty connections. The catch: this is a developer API, so there's no polished mobile app the way Otter.ai or Google's consumer apps feel. But for building, not just recording, that's the right call.

Daily Polish7.9

A full-featured free tier and clean single API feel considered.

Learning Curve7.7

One API for real-time and batch, but it still expects developers.

Mobile Parity7.5

A developer API with no consumer mobile app, neutral for this category.

Onboarding Experience8.3

8 free hours with no credit card lowers the starting bar.

Reliability Feel8.0

On-device option and a 20-year engine suggest dependable uptime.

Pros

  • Free tier is 8 hours monthly with full features and no card.
  • Code-switching follows speakers across languages naturally.
  • On-device deployment helps reliability on poor connections.
  • Single clean API for both live and recorded audio.

Cons

  • No consumer mobile app, since it's a developer API.
  • Still needs developer skills to get value out of it.

Right for

Developers who want a generous free tier to prototype voice features.

Avoid if

Non-technical users who want a ready-made transcription app.

The Skeptic

The Skeptic

Contrarian. Watch-outs, deal-breakers, broken promises, category patterns
7.6/10

Whisper made transcription free, so Speechmatics has to justify every dollar on accuracy.

In a category Whisper commoditized, Speechmatics earns its price on hard-audio accuracy, on-prem deployment, and a 2006 track record. The superlative marketing and stale 2022 funding are yellow flags, but the exit story and longevity are real.

The scariest word in speech-to-text right now is 'free.' OpenAI Whisper commoditized decent transcription overnight. So the real question is why anyone pays.

Speechmatics has an answer, and it's not marketing. Founded 2006, still independent while Nuance got swallowed by Microsoft. The Ursa engine's accuracy edge is real on hard audio — noisy, accented, multilingual — where Whisper drifts. On-prem and on-device deployment gives a clean exit story: your data never had to leave.

'The most accurate' is the kind of superlative that ages poorly, and there's no public per-language benchmark to check it. And the last funding round was 2022's $62M Series B — no fresh capital signal since. But 20 years and real enterprise use is a track record most rivals can't fake.

Competitive Differentiation7.3

Accuracy and on-device edge Whisper, but Deepgram contests the space hard.

Exit Portability8.0

Standard transcript output plus on-prem deployment ease any exit.

Long-term Viability7.2

Two decades independent, though no funding since the 2022 Series B.

Marketing Honesty7.0

The 'most accurate' superlative lacks a public per-language benchmark to verify.

Track Record Match8.0

A 2006 founding and enterprise use back the reliability claims.

Pros

  • Independent since 2006 while rivals like Nuance got acquired.
  • Accuracy edge on noisy and multilingual audio is real.
  • On-prem and on-device deployment gives a clean data-exit story.

Cons

  • The 'most accurate' claim has no public per-language benchmark.
  • No funding round since the 2022 Series B.

Right for

Buyers who need accuracy on noisy or multilingual audio.

Avoid if

Teams whose audio is clean enough for free Whisper.

Buyer Questions

Common questions answered by our AI research team

Pricing

Does Speechmatics have a free plan?

Yes. The Free tier gives 8 hours of speech-to-text per month (480 minutes), 2 concurrent real-time sessions, and 1,000,000 text-to-speech characters, with full API access and no credit card required.

Pricing

How much does Speechmatics cost per hour?

The Pro tier is usage-based, starting from $0.24 per hour of audio and billed to the second. Text-to-speech runs $0.011 per 1,000 characters, and volume discounts kick in above 500 hours per month.

Features

How many languages does Speechmatics support?

Speechmatics transcribes 56+ languages and translates speech to and from English across 34 languages. The Ursa engine also handles code-switching, so speakers can move between languages in the same audio.

Integration

Can Speechmatics build voice agents?

Yes, through the Flow API, which combines real-time ASR, large language models and text-to-speech into one conversational voice-agent solution. It integrates with real-time frameworks like LiveKit and Pipecat.

Setup

Can Speechmatics run on-premises?

Yes. Beyond the cloud API, Speechmatics deploys on-premises and on-device, suiting data-sensitive workloads in healthcare, finance and contact centers. Built-in speaker diarization and real-time translation work across all deployments.

Also in AI Voice & Speech