Broad Multilingual Reference

Self-Hosted Whisper vs API

Building a self-hosted Whisper pipeline is justified for strict offline data control or continuous high-volume workloads, but only if your measured total hardware and engineering costs beat per-minute pricing. A hosted API fits better when audio volume fluctuates unpredictably, engineering capacity is constrained, or your product requires turnkey timestamps without server maintenance.

When Each Option Fits

When Should You Self-Host?

  • Workloads governed by internal data-handling policies requiring recordings to remain on offline, owner-operated infrastructure, without providing an inherent guarantee of regulatory compliance.
  • High-volume, continuous processing pipelines where high utilization achieves an overall cost advantage over API billing only when factoring in measured hardware, hosting infrastructure, and engineering labor.

When Should You Use an API?

  • Workloads with intermittent or unpredictable audio volume where maintaining idle compute hardware increases effective per-minute processing costs.
  • Teams that require turnkey recorded transcription with structured JSON, SRT, or VTT outputs and optional word or speaker labels without managing runtime infrastructure. Pipelines with broad language or translation requirements must verify that candidate APIs support their specific target languages and outputs.

What Does the Model Support?

Architecture and Checkpoints
Encoder-decoder Transformer architecture released by OpenAI. Checkpoints include large-v3 and turbo. OpenAI Whisper
Audio Windowing Mechanism
Processes audio in 30-second sliding windows using log-mel spectrogram features. The reference transcribe implementation provides built-in windowing logic, avoiding the need to construct custom sliding-window mechanisms. OpenAI Whisper
Word-Level Alignment
Word timestamps are generated in reference transcribe.py via cross-attention weight matrices and dynamic time warping, adding computational passes over generated tokens. Whisper Transcription and Word Timing
Language Coverage and Translation
Whisper large-v3 supports multilingual transcription across diverse languages and speech-to-text translation into English. The turbo checkpoint focuses exclusively on multilingual transcription. OpenAI Whisper
Speaker Diarization
Whisper transcribes speech without native speaker separation. Attributing spoken segments requires an auxiliary pipeline such as pyannote Community 1, which assigns anonymous speaker labels rather than identities. pyannote Community-1 Diarization
Licensing
Source code and model weights are released under the MIT license. OpenAI Whisper

What You Need to Build

  1. Environment Isolation and Runtime Setup

    Deploy in an isolated container with pinned PyTorch and CUDA dependencies. Whisper can run on either GPU or CPU hardware. Pin the exact repository commit and checkpoint weights to maintain consistent decoding behavior.

    OpenAI Whisper
  2. Audio Standardization and Channel Handling

    Convert incoming audio to 16 kHz single-channel floating-point arrays. When source recordings contain dedicated speaker channels, preserve distinct channels before model-specific resampling rather than universally downmixing to mono.

    OpenAI Whisper
  3. Window Processing via Reference Code

    Call the supplied transcribe routine rather than implementing custom sliding-window code. The reference implementation already manages 30-second window offsets, padding, and temperature fallback schedules.

    OpenAI Whisper
  4. Word Timestamp Extraction

    Enable word_timestamps in the reference transcribe call to compute dynamic time warping alignment from cross-attention weights. Measure the added latency, as alignment requires additional computation per token.

    Whisper Transcription and Word Timing
  5. Auxiliary Diarization Integration

    If speaker separation is necessary, integrate pyannote Community 1 (weights under CC BY 4.0, code under MIT). Users must accept repository access conditions before downloading weights. Run the diarization pipeline on CPU or GPU, then align resulting anonymous speaker intervals with transcript words. When multi-track audio already isolates speakers on separate channels, map channels directly to avoid diarization.

    pyannote Community-1 Diarization

Operating a Transcription Service

Job Ingestion and Queue Management

Deploy durable job queues with distributed workers to ingest audio, isolate distinct speaker channels, enforce strict retry limits with idempotency keys, and emit reliable status callbacks to client applications upon job completion or terminal timeout.

Result Formatting and Storage Lifecycle

Export structured transcripts into JSON, SRT, or VTT formats by reconciling chunk offsets, aligned word boundaries, and speaker turn intervals, while enforcing retention rules to store intermediate payloads securely and promptly delete raw source recordings.

Production Observability and Quality Benchmarking

Monitor completed job latency, compute expense, and GPU memory failures in production, evaluate word error rates, speaker attribution, and timing across representative languages and long silences, and pin dependency upgrades to prevent unexpected decoding regressions.

Model tuning, data annotation, dedicated staff, and compliance programs are optional enterprise investments. They are not minimum self-hosting requirements for teams running their own speech recognition models in production environments.

How Much Hardware Do You Need?

OpenAI README documentation lists approximate inference VRAM of 10 GB for large-v3 and 6 GB for turbo, rather than bare weight loading footprints. Whisper can run on GPU or CPU hardware. To size production capacity, profile your pipeline with representative long audio files and every planned output stage. Test whether cross-attention word timestamps or auxiliary pyannote diarization bottleneck throughput. Because audio with extended silence or background noise can trigger repetition loops or decoding failures, benchmark decoding behavior across diverse recordings rather than assuming temperature fallback schedules prevent errors. Auxiliary pyannote diarization introduces separate compute overhead on CPU or GPU and assigns anonymous speaker labels. When recordings isolate speakers on separate channels, preserving multichannel audio avoids diarization entirely.

Whisper Model Sizes and Hardware

OpenAI Whisper ; Whisper Transcription and Word Timing ; pyannote Community-1 Diarization

Self-Hosting Cost Calculator

Defaults are illustrative assumptions shared across all four model guides, not measured model performance or hardware quotes. Replace throughput and costs with your own full-pipeline inputs.

Pipeline and Volume Settings
Infrastructure and Hardware Costs
Engineering Labor and Maintenance

Measure full-pipeline throughput including decoding and audio processing. Adjust throughput if adding unmeasured stages, though native timestamps may already be included.

Self-Hosted Estimate
$292.08
Speech is Cheap (Integration Included)
$51.25
Self-Hosted Minus API
$240.83
Compute and Hardware Cost
$36.00
Self-Hosted Extra CPU, Queues, and Monitoring
$31.08
Engineering Labor Cost
$225.00
API Service Fees
$20.00
API Integration and Maintenance Labor
$31.25
Active Processing GPU Hours
18
Maximum Monthly Capacity
--
Effective Cost per Minute
$0.013522
Projected Break-Even Volume
No break-even within 100 million monthly minutes or hardware capacity limits. Later cost crossings may still exist.

Break-even identifies the first positive whole minute where total self-hosted expenses are less than or equal to API billing when factoring in engineering effort on both sides. Crossing break-even does not guarantee self-hosting remains cheaper at higher volumes.

Owned hardware capacity is bounded by the configured cluster size, whereas cloud elastic provisioning assumes capacity dynamically matches audio volume.

This illustrative cost breakdown updates from current inputs to compare self-hosting expenses against Speech is Cheap pricing.

How Do Monthly Costs Compare?
Monthly Audio Volume (Minutes)Cloud Elastic HostingOwned Hardware FleetSpeech is Cheap (Integration Included)
600$256.03$368.36$32.45
6,000$265.30$368.63$43.25
21,600$292.08$369.41$51.25
60,000$358.00$371.33$86.81
600,000$1,285.00--$586.85

Evaluate projected monthly expenditure across 25%, 50%, and 80% GPU capacity utilization at your chosen audio volume, holding all other current pipeline and labor inputs constant.

Monthly Cost Sensitivity Across Utilization Levels
GPU UtilizationCloud Elastic HostingOwned Hardware FleetSpeech is Cheap (Integration Included)
25%$328.08$369.41$51.25
50%$292.08$369.41$51.25
80%$278.58$369.41$51.25

All monetary values are in USD. Audio volume is measured in whole minutes. A dash (--) indicates unavailable data, exceeded capacity, or not applicable.

Frequently Asked Questions

What Hardware Is Required to Run Whisper in Production?

The README lists approximate inference VRAM of 10 GB for large-v3 and 6 GB for turbo, though CPU execution is possible. Sizing requires testing your longest audio files and output stages like timestamps or diarization under concurrent load.

When Does Whisper Self-Hosting Cost Less Than an API?

Self-hosting costs less only when steady, high-volume audio maximizes GPU utilization enough to offset server hosting, queue infrastructure, and ongoing engineering maintenance. If audio traffic fluctuates or volume is low, pay-as-you-go API pricing yields lower total expenses.

Should You Build a Whisper Pipeline or Buy an API?

Build if you require strict data residency within private servers and have engineering staff to maintain decoding pipelines. Buy an API if you want instant timestamps, reliable scaling, and predictable per-minute costs without managing GPU clusters.

Where Are Specifications Sourced?

Choose Your Managed API Plan

Select from transparent monthly subscription or pay-as-you-go API plans from Speech is Cheap to support your production pipeline without the operational overhead of managing GPU infrastructure.

How Are Costs Calculated?

Default pipeline throughput of 1,200 audio minutes per active GPU hour (20x real-time) serves as a baseline planning assumption shared across all models, not an empirical benchmark of specific model or runtime performance.

Full pipeline throughput encompasses end-to-end processing, including audio decoding, CPU normalization, optional word timestamps, speaker diarization, and retry overhead. Certain models natively include word timestamps within primary decoding. All baseline cost figures are illustrative starting estimates and fully editable, rather than formal vendor quotes.

Cloud elastic hosting models compute expense by dividing monthly audio minutes by the product of pipeline throughput and GPU utilization, represented as M / (T * u), and multiplying by the hourly GPU rate. Billed time accounts for idle capacity once within that utilization factor, assuming elastic provisioning scales to meet incoming demand.

Owned hardware costs are calculated as N * (purchase / amortMonths + power) for N dedicated GPUs. Monthly processing capacity is modeled as 730 * N * u * T based on a standard 730-hour operating month. Standby capacity and reserve margins must be budgeted through utilization targets or additional hardware, as owned clusters offer no automatic redundancy.

Self-hosted infrastructure adds fixed monthly overhead for CPU ingestion workers, job queues, and observability, plus storage and egress calculated as M * storageRate. Shared application expenses outside speech decoding are equally excluded from both options. Engineering labor applies a single hourly rate to both approaches: self-hosted labor is laborRate * (setupHours / setupMonths + opsHours), while API labor is laborRate * (apiSetupHours / setupMonths + apiOpsHours). Setting labor hours to zero models incremental capacity on an existing team, not free engineering.

Speech is Cheap API expenses reflect the lowest available rate between pay-as-you-go and monthly subscription tiers from maintained pricing data, applying optional add-on fees across all processed minutes. In examples, API billing rounds each audio file up to the next whole minute and excludes taxes, volume credits, or custom negotiations. Projected break-even identifies the first positive whole minute where total self-hosted costs undercut API billing inclusive of engineering effort on both sides, bounded by 100 million minutes or owned hardware capacity; achieving break-even does not guarantee self-hosting remains cheaper at higher volumes.

How Is Capacity Calculated?

PAYGFees and subscriptionFees are the full plan bills including all selected add-ons; choose the lower amount rather than a price per minute. All other formulas use GPU hours and monthly minutes.

Cloud Compute Cost
cloudCost = (M / (T * u)) * gpuRate
Owned Hardware Cost
ownedCost = N * ((purchase / amortMonths) + power)
Owned Hardware Capacity
capacity = 730 * N * u * T
Self-Hosted Engineering Labor
selfLabor = laborRate * ((setupHours / setupMonths) + opsHours)
API Integration Labor
apiLabor = laborRate * ((apiSetupHours / setupMonths) + apiOpsHours)
Full Pipeline Totals
totalSelfHosted = computeCost + fixed + (M * storageRate) + selfLabor; totalAPI = min(PAYGFees, subscriptionFees) + apiLabor
M
Monthly audio minutes processed
T
Pipeline throughput in audio minutes per active GPU hour
u
Target GPU capacity utilization fraction (0 to 1.0)
N
Number of owned GPUs in cluster
gpuRate
Hourly cloud GPU rental rate in USD
purchase
Hardware purchase price per GPU in USD
amortMonths
Hardware depreciation amortization schedule in months
power
Monthly hosting, power, and cooling cost per GPU in USD
capacity
Maximum monthly audio processing capacity in minutes
fixed
Monthly fixed infrastructure cost for ingestion workers, queues, and monitoring in USD
storageRate
Storage and egress fee per audio minute in USD
laborRate
Blended engineering hourly rate in USD
setupHours
Initial engineering setup hours for self-hosting
setupMonths
Setup labor amortization period in months
opsHours
Ongoing monthly engineering maintenance hours for self-hosting
apiSetupHours
Initial engineering setup hours for API integration
apiOpsHours
Ongoing monthly engineering maintenance hours for API integration
computeCost
Monthly GPU compute expense for cloud elastic or owned hardware in USD
totalSelfHosted
Total monthly self-hosted pipeline cost in USD
totalAPI
Total monthly managed API cost in USD
Speech is Cheap Pricing ; Speech is Cheap Job Creation