Accuracy, Benchmarks, and Evaluation Guide

Speech‑to‑Text Accuracy

Speech is Cheap measured pipeline Word Error Rate on English LibriSpeech-Long audiobooks. Below, review our measured pipeline performance alongside separate published upstream benchmarks.

Normalized Word Error Rate

6.85%

95% Speaker-Bootstrap Interval

6.11% – 7.64%

Measured on eligible English audiobooks from LibriSpeech-Long across 476 recordings using our pinned inference pipeline.

Measured Pipeline Accuracy

Across the English LibriSpeech-Long corpus, Speech is Cheap evaluated 476 eligible audiobook recordings totaling 1,530.45 audio minutes. The pipeline achieved a 6.85% pooled normalized WER, with a 95% speaker-bootstrap confidence interval of 6.11% to 7.64%.

Recordings

476

Speakers

73

Audio Minutes

1,530.45

Evaluation Date

UTC

Speech is Cheap Pipeline Accuracy on LibriSpeech-Long Evaluated 2026-09-18 UTCPrimary English Normalizer
Corpus SplitRecordingsSpeakersReference WordsWord Error Rate (WER)95% Speaker-Bootstrap Interval
Test Clean26940146,3985.11%4.44% – 5.87%
Test Other20733106,4889.24%7.82% – 10.78%
Pooled47673252,8866.85%6.11% – 7.64%

This run recorded zero first-attempt failures across 476 recordings. Under our frozen scoring protocol for this evaluation, unreturned requests or timeouts are scored as empty hypotheses, not omitted. Word Error Rate measures edit operations against reference words, not percent correct.

Error Rate Confidence Intervals by Split

Displays Word Error Rate point estimates with 95% speaker confidence intervals by split, evaluated using the selected text normalization.

Test Clean 5.11% · 4.44% – 5.87%
Test Other 9.24% · 7.82% – 10.78%
Pooled 6.85% · 6.11% – 7.64%

Confidence intervals reflect 10,000 speaker-bootstrap draws, measuring statistical sensitivity to speaker mix in this corpus instead of universal expected error. They do not capture reference transcription errors, shared book context, domain bias, or pretraining exposure.

Recording Error Distribution

Distribution of per-recording Word Error Rates across selected splits and text normalization rules.

Recording Count

189
0–5
197
5–10
60
10–15
18
15–20
6
20–25
4
25–30
2
30–35
0
35+

Word Error Rate (%)

Selected Split Summary

Pooled · Primary English Normalizer

6.85%

95% Speaker-Bootstrap Interval 6.11% – 7.64%

17,319 Total Word Edits / 252,886 Reference Words

Edit Type Breakdown

Substitutions
6,930
Deletions
2,888
Insertions
7,501

test_other/6070-86745-0000

Word Error Rate (WER)
34.56%
Duration (Minutes)
4.00
Total Word Edits
197
Reference Words
570

Upstream English Model Benchmarks

The figures below come from the Hugging Face Open ASR Leaderboard, reporting upstream model evaluations across eight public English short-form datasets.

Public English Short-Form Datasets: Snapshot 2026-09-11
Evaluation DatasetNVIDIA Parakeet TDT 0.6B v3 WER, %OpenAI Whisper large-v3-turbo WER, %
AMI-Cleaned9.4213.88
Earnings22-Cleaned-AA-chunked5.858.09
Gigaspeech-Cleaned7.998.47
LS Clean1.522.13
LS Other3.133.71
SPGISpeech3.632.79
Voice Arena Monsoon4.144.77
Voxpopuli-AA-Cleaned3.197.02
Unweighted Mean of Eight Dataset WERs4.866.36

Values represent Word Error Rate (WER) percentages where lower is better. Parakeet TDT 0.6B v3 achieves a lower mean of 4.86% across the suite, while Whisper large-v3-turbo scores lower on SPGISpeech at 2.79% compared to 3.63%. Relative performance varies across audio domains.

Hugging Face: Pinned English Results CSV

Understanding Word Error Rate

Word Error Rate compares a reference transcript to model output using minimum word insertions, deletions, and substitutions divided by the reference word count. Lower scores are better. Because it measures an edit rate rather than probability or percent accuracy, WER can exceed 100% and depends on text normalization.

WER=S+D+IN×100%
S
Substitutions
D
Deletions
I
Insertions
N
Reference Words

Calculating WER

For an illustrative 100-word reference, 3 substitutions, 2 deletions, and 1 insertion total 6 edits, giving a 6% WER. The three small examples below are also invented scenarios rather than observed API results.

Substitution

Reference
send the invoice
Transcript
send an invoice

The model replaces a reference word with a different word, counting as one edit.

Deletion

Reference
please call tomorrow
Transcript
please call

The model leaves out a reference word, counting as one edit.

Insertion

Reference
call tomorrow
Transcript
please call tomorrow

The model adds an extra word not present in the reference transcript, counting as one edit.

What WER Leaves Out

A low benchmark score does not guarantee good results in real-world applications. Word Error Rate measures word matching across a test set, but it overlooks practical system requirements. A low average score is not a reliability metric, and failed jobs require separate tracking.

Normalization Rules and Text Changes

Normalized lexical scoring can remove punctuation and case before counting word edits. Assess readable formatting separately from word accuracy.

Short Clips and Long Recordings

An aggregate score across English short-form benchmarks does not establish performance on your long recordings, other languages, background noise or overlapping speech. Test these conditions on your own recordings.

Names, Numbers and Context

Standard Word Error Rate treats all word edits equally. However, incorrect names, numbers, and key terms often matter much more to your business than minor words. Track critical terms and identifying numbers alongside general error rates.

Average Scores and System Failures

Word matching does not measure speaker attribution, timestamp precision or service reliability. Record timeouts, retries and empty outputs separately using a predefined failure policy. An empty transcript of silent audio is different from a failed request.

How to Evaluate Accuracy on Your Audio

To determine which transcription API fits your workload, evaluate performance on your own audio rather than relying on vendor benchmarks. Follow this four-step evaluation process to establish reproducible results.

  1. Freeze a Representative Corpus

    Select rights-cleared representative recordings and human-verified references before seeing any model results. Record audio minutes, languages, domain, background noise, speaker count, audio channels, sample rates, and explicit exclusions. Freezing these assets in advance keeps your evaluations consistent over time.

  2. Standardize Configurations and Normalization

    Record run dates, API options, and model versions when available, while noting unknown provider revisions. Freeze text normalization rules for punctuation, case, numbers, contractions, fillers, and non-speech sounds. Set clear scoring rules in advance for failed, timed-out, retried, and empty outputs.

  3. Score Aligned Segments and Record Failures

    Score assembled transcripts in original chunk order along correct reference boundaries. Calculate per-condition WER with sample and reference-word counts, keeping pooled WER distinct from dataset means. Estimate uncertainty where sample size allows, and ensure failed jobs are reported rather than dropped.

  4. Check More Than Words

    Assess punctuation, speaker attribution, and timestamp alignment separately from lexical word error rates. Track request latency and transcription cost alongside error metrics. Evaluating these operational factors alongside accuracy ensures you choose the right API for your workload.

Test Your Audio on Speech is Cheap

Evaluate our transcription API using your own recordings. Test transcription on our free demo, or choose an API subscription or pay-as-you-go plan tailored to your volume.

Measurement Methodology and Technical Scope

Corpus Selection and Predeclared Exclusions

We evaluated the LibriSpeech-Long English audiobook corpus (CC BY 4.0, 16 kHz mono). Full dataset revisions, commit pins, and seeds are recorded in our downloadable verification files. The eligible corpus comprises 476 recordings totaling 1,530.45 minutes across 73 speakers, with recordings up to 4 minutes long. One predeclared 5.34-second clip below the six-second minimum was excluded. Twelve development pilot recordings and one timing repeat were also excluded. The corpus was frozen before execution.

Pipeline Execution and Request Settings

Inference was executed on 2026-09-18 UTC using the same pinned inference image as the API. Audio was processed in 30-second segments using automatic language detection and a 0.5 confidence threshold. Optional features like timestamps and diarization were disabled. The 469 segments flagged non-English were retained as generated model output, not verified linguistic labels. Periodic operational checks monitored pipeline health without inspecting accuracy scores.

Text Normalization and Alignment Rules

Primary scoring uses the unchanged OpenAI Whisper English normalizer to standardize casing, punctuation, numbers, contractions, and fillers. Alignment uses jiwer and RapidFuzz. Word Error Rate is computed by dividing summed substitutions, deletions, and insertions by total normalized reference words across each split, without averaging recording percentages or manual transcript correction. Secondary lexical scoring evaluates identical predictions. Software versions and verification hashes are listed in downloadable assets.

Bootstrap Resampling and Uncertainty Estimation

Confidence intervals derive from 10,000 bootstrap draws with speaker resampling within each split, retaining all recordings per chosen speaker. Recalculating word-weighted error ratios yields the 2.5th and 97.5th percentiles for the 95% interval. This estimates sensitivity to speaker mix in this corpus, not universal expected error. It does not model shared literary context, transcription errors, domain sampling bias, or training exposure. Resampling parameters are preserved in downloadable assets.

Scope Boundaries and Benchmark Limits

This evaluation characterizes inference on English read audiobooks with recordings up to 4 minutes long. These findings do not cover conversational dialogue, overlapping speech, non-English speech, diarization, or cloud API throughput. Word Error Rate reflects lexical string alignment, not punctuation or semantic meaning.

Verification Data and Reproduction Bundle

Offline verification assets include summary JSON, per-recording metrics CSV, predictions JSON, manifest hashes, and Docker reproduction scripts. You can independently verify hashes, recompute normalized edit counts, and reproduce bootstrap intervals offline without downloading raw audio. An interactive recording explorer is hosted on our CDN.

LibriSpeech-Long dataset provided under CC BY 4.0 (Park et al., Long-Form Speech Generation with Spoken Language Models, 2024; Panayotov et al., LibriSpeech: An ASR Corpus Based on Public Domain Audio Books, 2015). The OpenAI Whisper text normalizer is distributed under the MIT License.

Speech is Cheap: Pipeline Results ; Speech is Cheap: Evaluation Methodology ; Speech is Cheap: Offline Reproduction Bundle ; LibriSpeech-Long: Pinned Dataset

Scoring Notes and Benchmark History

Acoustic conditions like background noise, microphone quality, and speaker clarity directly affect transcription accuracy. Evaluate the API with your own recordings using the interactive demo.

This table includes all eight public datasets and excludes the two private aggregates. As a result, this eight-dataset mean is not the leaderboard's default overall score or an official rank. The summary row reflects the unweighted macro-mean of all eight public datasets, not an error rate pooled across words, customers, or minutes.

Hugging Face: Pinned English Results CSV ; Hugging Face: Leaderboard Snapshot Registry ; Hugging Face: Private Evaluation Data

The published-scores download reproduces verified source extractions and dataset means from published leaderboard records. It does not reflect model inference runs or measured Speech is Cheap results. The CSV omits each model's run date, runtime revision, artifact revision, decoding settings, audio sample count, reference-word count, and confidence interval. These fields are null in the download.

At the inspected revision, current repository code resamples to 16 kHz, applies EnglishTextNormalizer, and filters empty or ignored references. It also uses compound merge scoring and session assembly. However, this code does not prove the exact historical execution configuration used for earlier benchmark runs.

Hugging Face: Reference Preparation and Normalization ; Hugging Face: WER Scoring and Transcript Assembly

The leaderboard snapshot is dated 2026-09-11. We checked the source extraction and arithmetic on 2026-09-17 UTC. Neither date identifies the original model evaluation runs.

Hugging Face: Pinned English Results CSV ; Hugging Face: Leaderboard Snapshot Registry

The current NVIDIA model card reports a historical 6.34% error rate on an older suite including TED-LIUM v3 instead of Voice Arena Monsoon. Because the evaluation suite changed, the difference from 4.86% is not evidence of model or API improvement.

NVIDIA Parakeet TDT 0.6B v3 Model Card

Multilingual model-card evaluation suites have different language coverage and test conditions. You cannot make universal inferences about other languages from English benchmark scores. Always evaluate transcription accuracy on your specific target languages directly.

NVIDIA Parakeet TDT 0.6B v3 Model Card