Substitution
- Reference
- send the invoice
- Transcript
- send an invoice
The model replaces a reference word with a different word, counting as one edit.
Accuracy, Benchmarks, and Evaluation Guide
Speech is Cheap measured pipeline Word Error Rate on English LibriSpeech-Long audiobooks. Below, review our measured pipeline performance alongside separate published upstream benchmarks.
6.85%
95% Speaker-Bootstrap Interval
6.11% – 7.64%
Measured on eligible English audiobooks from LibriSpeech-Long across 476 recordings using our pinned inference pipeline.
Across the English LibriSpeech-Long corpus, Speech is Cheap evaluated 476 eligible audiobook recordings totaling 1,530.45 audio minutes. The pipeline achieved a 6.85% pooled normalized WER, with a 95% speaker-bootstrap confidence interval of 6.11% to 7.64%.
Recordings
476
Speakers
73
Audio Minutes
1,530.45
Evaluation Date
UTC
Our primary normalizer applies the OpenAI Whisper English normalizer, summing edit operations across normalized reference words without manual correction. A secondary lexical scorer standardizes only Unicode, casing, punctuation, and whitespace (7.17% pooled). This sensitivity check evaluates the identical saved predictions without claiming parity with external leaderboards.
| Corpus Split | Recordings | Speakers | Reference Words | Word Error Rate (WER) | 95% Speaker-Bootstrap Interval |
|---|---|---|---|---|---|
| Test Clean | 269 | 40 | 146,398 | 5.11% | 4.44% – 5.87% |
| Test Other | 207 | 33 | 106,488 | 9.24% | 7.82% – 10.78% |
| Pooled | 476 | 73 | 252,886 | 6.85% | 6.11% – 7.64% |
This run recorded zero first-attempt failures across 476 recordings. Under our frozen scoring protocol for this evaluation, unreturned requests or timeouts are scored as empty hypotheses, not omitted. Word Error Rate measures edit operations against reference words, not percent correct.
Displays Word Error Rate point estimates with 95% speaker confidence intervals by split, evaluated using the selected text normalization.
Confidence intervals reflect 10,000 speaker-bootstrap draws, measuring statistical sensitivity to speaker mix in this corpus instead of universal expected error. They do not capture reference transcription errors, shared book context, domain bias, or pretraining exposure.
Distribution of per-recording Word Error Rates across selected splits and text normalization rules.
The split filter updates the distribution chart, selected split summary, and recording selector, while the three-split comparison remains visible.
Recording Count
Word Error Rate (%)
Pooled · Primary English Normalizer
6.85%
95% Speaker-Bootstrap Interval 6.11% – 7.64%
17,319 Total Word Edits / 252,886 Reference Words
Select a public recording identifier to view audio duration, edit counts, and Word Error Rate.
test_other/6070-86745-0000
The figures below come from the Hugging Face Open ASR Leaderboard, reporting upstream model evaluations across eight public English short-form datasets.
| Evaluation Dataset | NVIDIA Parakeet TDT 0.6B v3 WER, % | OpenAI Whisper large-v3-turbo WER, % |
|---|---|---|
| AMI-Cleaned | 9.42 | 13.88 |
| Earnings22-Cleaned-AA-chunked | 5.85 | 8.09 |
| Gigaspeech-Cleaned | 7.99 | 8.47 |
| LS Clean | 1.52 | 2.13 |
| LS Other | 3.13 | 3.71 |
| SPGISpeech | 3.63 | 2.79 |
| Voice Arena Monsoon | 4.14 | 4.77 |
| Voxpopuli-AA-Cleaned | 3.19 | 7.02 |
| Unweighted Mean of Eight Dataset WERs | 4.86 | 6.36 |
Values represent Word Error Rate (WER) percentages where lower is better. Parakeet TDT 0.6B v3 achieves a lower mean of 4.86% across the suite, while Whisper large-v3-turbo scores lower on SPGISpeech at 2.79% compared to 3.63%. Relative performance varies across audio domains.
Hugging Face: Pinned English Results CSVWord Error Rate compares a reference transcript to model output using minimum word insertions, deletions, and substitutions divided by the reference word count. Lower scores are better. Because it measures an edit rate rather than probability or percent accuracy, WER can exceed 100% and depends on text normalization.
For an illustrative 100-word reference, 3 substitutions, 2 deletions, and 1 insertion total 6 edits, giving a 6% WER. The three small examples below are also invented scenarios rather than observed API results.
The model replaces a reference word with a different word, counting as one edit.
The model leaves out a reference word, counting as one edit.
The model adds an extra word not present in the reference transcript, counting as one edit.
A low benchmark score does not guarantee good results in real-world applications. Word Error Rate measures word matching across a test set, but it overlooks practical system requirements. A low average score is not a reliability metric, and failed jobs require separate tracking.
Normalized lexical scoring can remove punctuation and case before counting word edits. Assess readable formatting separately from word accuracy.
An aggregate score across English short-form benchmarks does not establish performance on your long recordings, other languages, background noise or overlapping speech. Test these conditions on your own recordings.
Standard Word Error Rate treats all word edits equally. However, incorrect names, numbers, and key terms often matter much more to your business than minor words. Track critical terms and identifying numbers alongside general error rates.
Word matching does not measure speaker attribution, timestamp precision or service reliability. Record timeouts, retries and empty outputs separately using a predefined failure policy. An empty transcript of silent audio is different from a failed request.
To determine which transcription API fits your workload, evaluate performance on your own audio rather than relying on vendor benchmarks. Follow this four-step evaluation process to establish reproducible results.
Select rights-cleared representative recordings and human-verified references before seeing any model results. Record audio minutes, languages, domain, background noise, speaker count, audio channels, sample rates, and explicit exclusions. Freezing these assets in advance keeps your evaluations consistent over time.
Record run dates, API options, and model versions when available, while noting unknown provider revisions. Freeze text normalization rules for punctuation, case, numbers, contractions, fillers, and non-speech sounds. Set clear scoring rules in advance for failed, timed-out, retried, and empty outputs.
Score assembled transcripts in original chunk order along correct reference boundaries. Calculate per-condition WER with sample and reference-word counts, keeping pooled WER distinct from dataset means. Estimate uncertainty where sample size allows, and ensure failed jobs are reported rather than dropped.
Assess punctuation, speaker attribution, and timestamp alignment separately from lexical word error rates. Track request latency and transcription cost alongside error metrics. Evaluating these operational factors alongside accuracy ensures you choose the right API for your workload.
Checked
Checked
Checked
Checked
Evaluate our transcription API using your own recordings. Test transcription on our free demo, or choose an API subscription or pay-as-you-go plan tailored to your volume.
We evaluated the LibriSpeech-Long English audiobook corpus (CC BY 4.0, 16 kHz mono). Full dataset revisions, commit pins, and seeds are recorded in our downloadable verification files. The eligible corpus comprises 476 recordings totaling 1,530.45 minutes across 73 speakers, with recordings up to 4 minutes long. One predeclared 5.34-second clip below the six-second minimum was excluded. Twelve development pilot recordings and one timing repeat were also excluded. The corpus was frozen before execution.
Inference was executed on 2026-09-18 UTC using the same pinned inference image as the API. Audio was processed in 30-second segments using automatic language detection and a 0.5 confidence threshold. Optional features like timestamps and diarization were disabled. The 469 segments flagged non-English were retained as generated model output, not verified linguistic labels. Periodic operational checks monitored pipeline health without inspecting accuracy scores.
Primary scoring uses the unchanged OpenAI Whisper English normalizer to standardize casing, punctuation, numbers, contractions, and fillers. Alignment uses jiwer and RapidFuzz. Word Error Rate is computed by dividing summed substitutions, deletions, and insertions by total normalized reference words across each split, without averaging recording percentages or manual transcript correction. Secondary lexical scoring evaluates identical predictions. Software versions and verification hashes are listed in downloadable assets.
Confidence intervals derive from 10,000 bootstrap draws with speaker resampling within each split, retaining all recordings per chosen speaker. Recalculating word-weighted error ratios yields the 2.5th and 97.5th percentiles for the 95% interval. This estimates sensitivity to speaker mix in this corpus, not universal expected error. It does not model shared literary context, transcription errors, domain sampling bias, or training exposure. Resampling parameters are preserved in downloadable assets.
This evaluation characterizes inference on English read audiobooks with recordings up to 4 minutes long. These findings do not cover conversational dialogue, overlapping speech, non-English speech, diarization, or cloud API throughput. Word Error Rate reflects lexical string alignment, not punctuation or semantic meaning.
Offline verification assets include summary JSON, per-recording metrics CSV, predictions JSON, manifest hashes, and Docker reproduction scripts. You can independently verify hashes, recompute normalized edit counts, and reproduce bootstrap intervals offline without downloading raw audio. An interactive recording explorer is hosted on our CDN.
LibriSpeech-Long dataset provided under CC BY 4.0 (Park et al., Long-Form Speech Generation with Spoken Language Models, 2024; Panayotov et al., LibriSpeech: An ASR Corpus Based on Public Domain Audio Books, 2015). The OpenAI Whisper text normalizer is distributed under the MIT License.
Speech is Cheap: Pipeline Results ; Speech is Cheap: Evaluation Methodology ; Speech is Cheap: Offline Reproduction Bundle ; LibriSpeech-Long: Pinned DatasetAcoustic conditions like background noise, microphone quality, and speaker clarity directly affect transcription accuracy. Evaluate the API with your own recordings using the interactive demo.
This table includes all eight public datasets and excludes the two private aggregates. As a result, this eight-dataset mean is not the leaderboard's default overall score or an official rank. The summary row reflects the unweighted macro-mean of all eight public datasets, not an error rate pooled across words, customers, or minutes.
Hugging Face: Pinned English Results CSV ; Hugging Face: Leaderboard Snapshot Registry ; Hugging Face: Private Evaluation DataThe published-scores download reproduces verified source extractions and dataset means from published leaderboard records. It does not reflect model inference runs or measured Speech is Cheap results. The CSV omits each model's run date, runtime revision, artifact revision, decoding settings, audio sample count, reference-word count, and confidence interval. These fields are null in the download.
At the inspected revision, current repository code resamples to 16 kHz, applies EnglishTextNormalizer, and filters empty or ignored references. It also uses compound merge scoring and session assembly. However, this code does not prove the exact historical execution configuration used for earlier benchmark runs.
Hugging Face: Reference Preparation and Normalization ; Hugging Face: WER Scoring and Transcript AssemblyThe leaderboard snapshot is dated 2026-09-11. We checked the source extraction and arithmetic on 2026-09-17 UTC. Neither date identifies the original model evaluation runs.
Hugging Face: Pinned English Results CSV ; Hugging Face: Leaderboard Snapshot RegistryThe current NVIDIA model card reports a historical 6.34% error rate on an older suite including TED-LIUM v3 instead of Voice Arena Monsoon. Because the evaluation suite changed, the difference from 4.86% is not evidence of model or API improvement.
NVIDIA Parakeet TDT 0.6B v3 Model CardMultilingual model-card evaluation suites have different language coverage and test conditions. You cannot make universal inferences about other languages from English benchmark scores. Always evaluate transcription accuracy on your specific target languages directly.
NVIDIA Parakeet TDT 0.6B v3 Model Card