Speech-to-Text Shortlists

Best Open‑Source Speech‑to‑Text Models

Start with Whisper for a versatile multilingual baseline. Look at Parakeet for European-language word timestamps, Canary for two-way translation, Qwen3-ASR for Chinese dialects, or Vosk for local CPU setups. These five options solve different technical problems, and running downloadable weights requires you to manage your own transcription service.

How to Choose

Written for software engineers and systems developers evaluating recognition models to deploy, host, and run on their own local machines or server infrastructure.

We include four model families and one recognition toolkit with downloadable weights, official documentation, and clear deployment boundaries. Code and weights are labeled by license. Hosted cloud APIs are intentionally excluded from this list.

  • Target language and task requirements, including transcription versus spoken language translation.
  • Runtime environment, compute resources, and the operational effort required to serve production traffic.
  • Permissive code and weight licenses, timestamp support, and missing components like speaker separation.

Whisper appears first as the general multilingual baseline, followed by specialized options for word timing, translation, regional dialects, and offline CPU setups. The order reflects deployment focus, not a performance ranking.

Five Options for Self-Hosting

Review hardware requirements, supported languages, and missing pipeline components before benchmarking these open-source models on your own recorded audio.

Open-Source Model

OpenAI Whisper

Best For: A multilingual baseline across multiple model sizes

Start here for a proven multilingual baseline to test across different hardware setups.

Strengths
Whisper handles multilingual transcription, language identification, and English translation. Multiple checkpoint sizes let you balance memory needs against quality.
Limitations
Model quality varies by language, and the model can hallucinate. Whisper Turbo is not trained for translation. Speaker separation requires another component.
Product Boundary
Downloadable weights and inference code only. The hosted OpenAI API is a separate commercial product.
Cost and License
MIT license for code and weights. You supply compute and manage operations.
Compare Self-Hosting and an API

Open-Weight Model

NVIDIA Parakeet-TDT-0.6B-v3

Best For: European-language recordings that require word timestamps

Choose Parakeet when your audio is among its 25 supported European languages and output contracts demand timestamps.

Strengths
The model outputs capitalization, punctuation, word timestamps, and segment timestamps. Its local attention mechanism accommodates longer recordings than standard full attention.
Limitations
Supported languages are narrower than Whisper. Long-audio throughput depends on attention configuration and GPU hardware, while speaker diarization requires a separate tool.
Product Boundary
An open-weight model run through tools like NeMo, not an out-of-the-box managed transcription endpoint.
Cost and License
Model weights use the CC BY 4.0 attribution license. You fund compute and infrastructure.
Compare Self-Hosting and an API

Open-Weight Model

NVIDIA Canary-1B-v2

Best For: Bi-directional translation between English and European languages

Select Canary when your pipeline needs direct spoken translation alongside transcription across supported European languages.

Strengths
Canary transcribes 25 languages and translates between English and the other 24. Explicit control tokens let your application select transcription or translation tasks.
Limitations
Canary v2 supports native timestamps and long-form inference through its documented runtime. However, it does not support arbitrary language pair translation. Runtime and hardware need evaluation on long recordings. Speaker diarization requires a separate stage.
Product Boundary
A self-hosted speech model for translation and transcription. Spoken language translation differs from verbatim transcription.
Cost and License
Weights use CC BY 4.0. You must budget for hosting compute and operational maintenance.
Compare Self-Hosting and an API

Open-Source Model

Qwen3-ASR

Best For: Chinese dialects and workloads needing streaming or recorded input

Choose Qwen3-ASR when Chinese speech and regional dialects are central to your workload.

Strengths
Supports 30 languages, 22 Chinese dialects, and both streaming and recorded inference. Providing context can help with specialized terms.
Limitations
Word timestamps require an external aligner with its own language limits. Runtime streaming code lacks production queueing and speaker diarization.
Product Boundary
Downloadable weights and inference scripts. Production service management and audio alignment remain separate tasks.
Cost and License
Apache 2.0 license for code and models. You supply compute and operational management.
Compare Self-Hosting and an API

Open-Source Engine and Models

Vosk

Best For: Offline CPU applications with fixed language requirements

Use Vosk when building offline applications on modest hardware without cloud connectivity or dedicated accelerators.

Strengths
Features streaming recognition, compact language-specific models, and programming bindings for multiple languages. Runs on desktop CPUs, mobile hardware, and Raspberry Pi.
Limitations
Each language requires its own model with distinct quality tradeoffs. Model licenses vary, meaning the engine license does not cover every download.
Product Boundary
A local recognition toolkit and model collection, not a universal multilingual checkpoint or hosted service.
Cost and License
Engine code is Apache 2.0. Individual model licenses vary; the small US English model is Apache 2.0.

When a Hosted API Is the Better Fit

Use self-hosting for offline operation, model control, or your own infrastructure. Include preprocessing, speakers, retries, storage, and monitoring in your estimate.

Speech is Cheap offers a hosted alternative for recorded audio, so it sits outside this self-hosted list. Compare its paid plans when you want transcription without managing server infrastructure. Use our comparison guides below to evaluate that trade-off.

Sources and Verification

Written and Published by Speech is Cheap. Updated .

Speech is Cheap publishes this guide and sells a service discussed in it. Recommendations rely on workload criteria rather than composite scores.

Dates show when we verified each source, not original documentation publication dates. We review sources quarterly and recheck them before updates.

  1. Canary-1B-v2 Model Card

    Checked

  2. NVIDIA NeMo Toolkit

    Checked

  3. OpenAI Whisper Repository

    Checked

  4. Parakeet-TDT-0.6B-v3 Model Card

    Checked

  1. Qwen3 Forced Aligner

    Checked

  2. Qwen3-ASR Repository

    Checked

  3. Qwen3-ASR-0.6B Model Card

    Checked

  4. Qwen3-ASR-1.7B Model Card

    Checked

  1. Vosk Models and Licenses

    Checked

  2. Vosk Offline Recognition

    Checked

  3. Whisper Model Card

    Checked

Choose a Plan for Your Recordings

Review included minutes, overage rates and optional features before you integrate. Pick the Speech is Cheap subscription or pay-as-you-go plan that matches your monthly recorded volume.

Editorial and Billing Notes

License labels reflect official project repositories. Always verify specific model weights, runtimes, and dependencies before deployment, because a permissive software license does not automatically extend to model weights or training data.

Published benchmark numbers rely on differing test sets, hardware, and evaluation methodologies. We have not conducted head-to-head testing for this article, and we make no claims regarding speed or accuracy winners.