Speech-to-Text Shortlists

Best Speech‑to‑Text APIs

Select Deepgram for live and recorded audio, AssemblyAI for transcript analysis, Google Cloud for deferred batch jobs, Speechmatics for audio events, OpenAI for existing platform integrations, or Speech is Cheap for recurring recorded volume. Determine your required endpoints, features, and output formats before comparing rates across providers.

How to Choose

Designed for developers and technical teams selecting a hosted speech-to-text API for software applications, rather than end users looking for a consumer dictation app.

These six services represent distinct developer use cases with published API documentation and pricing. Other providers appear in our broader pricing comparison. Their absence here is not an editorial slight; this is a focused shortlist, not an exhaustive market ranking.

  • Audio input mode: whether your application processes live streams or submitted audio recordings.
  • Output contracts: support for required languages, file size caps, speaker labels, and word timestamps.
  • Effective monthly cost: base transcription rates, optional add-on fees, and minimum commitment levels.

The list progresses through live audio, transcript analysis, deferred batch, audio events, existing vendor SDKs, and recurring recorded volume. Each entry addresses a distinct operational requirement. We name no overall winner.

Six APIs by Workload

The rate shown for each option reflects base recorded transcription for a representative model. Real-time streaming endpoints and optional add-on features will adjust your final bill.

Hosted API

Deepgram

Best For: Applications requiring both live streaming and recorded audio transcription

Shortlist Deepgram when your product roadmap combines real-time conversational audio and recorded batch files under one vendor account.

Strengths
Nova-3 pricing for recorded audio includes word timestamps and speaker diarization. Dedicated streaming endpoints support interactive, low-turnaround voice workflows.
Limitations
Pricing and billing rules differ between recorded and live endpoints. Language availability varies across specific models and processing modes.
Product Boundary
Recorded REST endpoints and live streaming WebSockets are distinct integrations. Quoted file transcription rates do not apply to live voice agents.
Cost and License
Nova-3 Monolingual PAYG: $0.0043 per minute.
Compare Deepgram for Recorded Audio

Hosted API

AssemblyAI

Best For: Workloads that require transcription alongside speech understanding capabilities

Evaluate AssemblyAI when your product pipeline needs managed speech understanding features and prompting built directly into the transcription workflow.

Strengths
Base recorded transcription provides word timing by default. The platform offers broad speech understanding tools, LLM prompting, and a separate streaming API.
Limitations
Speaker diarization and audio intelligence features incur additional charges. Universal-2 supports 99 languages, whereas Universal-3.5 Pro covers only 18.
Product Boundary
Recorded transcription, streaming endpoints, and speech understanding tools carry distinct pricing structures and feature sets.
Cost and License
Universal-2: $0.0025 per minute.
Compare AssemblyAI for Recorded Audio

Hosted Cloud API

Google Cloud Speech-to-Text

Best For: Cloud Storage workflows that can tolerate deferred batch processing

Consider Google Cloud when your audio files already reside in Cloud Storage buckets and immediate turnaround is unnecessary.

Strengths
Dynamic batch processing delivers lower per-minute rates than standard recognition. The Chirp 3 model supports speech adaptation, diarization, and regional deployment.
Limitations
Feature support depends on the chosen region, language, and method. Chirp 3 limits audio to 60 minutes for batch jobs, or 20 minutes when requesting word timestamps.
Product Boundary
This recommendation covers Speech-to-Text V2. Dynamic batch is deferred processing, and V1 free allowances do not apply to V2.
Cost and License
Dynamic Batch, Chirp 3: $0.003 per minute.
Compare Google Cloud Batch Options

Hosted API with Deployment Options

Speechmatics

Best For: Audio files requiring timestamps, diarization, and non-speech event detection

Choose Speechmatics Standard or Enhanced models when your application needs transcriptions that capture non-speech sounds alongside spoken words.

Strengths
Standard and Enhanced include word timing, speaker diarization, and Audio Events detection. Speechmatics also provides on-premises deployment options.
Limitations
The lower-priced Melia 1 model does not support Audio Events. You must verify feature availability across deployment modes and separate audio events from spoken dialogue.
Product Boundary
Cloud batch, real-time streaming, and on-premises software are distinct products. Selecting a lower-priced tier means forfeiting certain feature capabilities.
Cost and License
Standard: $0.004 per minute.
Compare Speechmatics Models and Costs

Hosted API

OpenAI

Best For: Recorded file transcription within an existing OpenAI infrastructure setup

Select OpenAI when your application already uses OpenAI SDKs and requires straightforward recorded audio transcription under existing billing.

Strengths
The current documentation recommends GPT-Transcribe, supporting custom context, prompt keywords, and language hints. It can stream text output as completed files process.
Limitations
Uploads are restricted to roughly 0.023 GiB per file. Legacy Whisper and listed GPT-4o transcription models will retire on 2027-02-26.
Product Boundary
Hosted GPT-Transcribe, the legacy hosted Whisper endpoint, and open-source Whisper are separate offerings. Live microphone audio requires a separate real-time API.
Cost and License
GPT-Transcribe: $0.0045 per minute.
Compare OpenAI Transcription Options

Hosted Recorded-Audio API

Speech is Cheap

Choose an API Plan

Best For: Recurring recorded audio volume and long single-file jobs

Include us when you process recordings regularly and want a monthly allowance rather than standalone per-minute pricing.

Strengths
Handles audio up to 1,440 minutes and direct uploads under 2 GiB. Supports optional speaker labels, word timestamps, sound events, and subtitles.
Limitations
No live streaming or self-hosted option. Optional add-ons cost extra, and subscriptions do not suit occasional small tasks.
Product Boundary
A dedicated paid API for recorded audio. The public browser demo does not supply free API access.
Cost and License
$20 per month includes 21,600 base minutes; overage $0.000926 per minute. PAYG: $0.002 per minute.
Review Speech is Cheap Plans

Test the Output Your Application Uses

Benchmark your short-listed providers against an identical set of realistic audio recordings. Evaluate transcription errors, domain vocabulary, regional accents, background noise, and timestamp alignment against ground truth. Track queuing delay separately from raw model processing time.

If building interactive products, evaluate live streaming endpoints and conversational turn detection. For recorded audio archives, test end-to-end batch workflows, including error handling and retry logic. Clean text alone is insufficient if speaker labels or timestamp drift break your application.

Sources and Verification

Written and Published by Speech is Cheap. Updated .

Speech is Cheap publishes this guide and sells a service discussed in it. Recommendations rely on workload criteria rather than composite scores.

Dates show when we verified each source, not original documentation publication dates. We review sources quarterly and recheck them before updates.

Choose a Plan for Your Recordings

Review included minutes, overage rates and optional features before you integrate. Pick the Speech is Cheap subscription or pay-as-you-go plan that matches your monthly recorded volume.

Editorial and Billing Notes

Prices reflect USD base rates for recorded audio transcription. Figures are sourced from our shared pricing comparison, excluding taxes, promotional credits, volume discounts, and add-on charges. Hourly vendor fees are normalized to minutes; tier rules remain in the full comparison.

OpenAI publishes a 25 MB file upload cap, which corresponds to roughly 0.023 GiB using decimal conversion. One GiB equals 1,073,741,824 bytes.

These editorial recommendations reflect published documentation rather than controlled benchmark testing. Always verify current language coverage, model versions, and regional hosting availability before finalizing your production integration.