OpenAI Whisper Explained

What Is OpenAI Whisper?

Whisper is an open-source speech recognition model developed by OpenAI that converts recorded audio into text. Released under the MIT license, it performs multilingual transcription, language identification, and speech translation into English. Whisper processes recorded audio files rather than operating as an interactive chatbot, text-to-speech engine, or turnkey cloud service.

What Does Whisper Include?

Automatic speech recognition converts spoken audio into written text. Whisper performs three core tasks: transcribing speech in its original language, translating non-English speech directly into English text, and identifying the spoken language in a recording. While trained on diverse audio, transcription accuracy varies across speakers, background noise, recording equipment, and technical terminology.

Evaluating Whisper requires distinguishing the trained neural network from the software used to run it. Developers separate model weights from local runtime libraries, self-hosted servers, and commercial cloud APIs. Choosing an approach depends on data privacy needs, technical resources, and budget.

Model Weights and Checkpoints
Trained neural network parameters released under the MIT license across multiple sizes, converting speech patterns into written text.
Reference Software Implementation
The official OpenAI Python package providing command-line tools and PyTorch code to run inference on CPU or GPU hardware.
Self-Hosted Production Stack
The surrounding software required to run Whisper reliably at scale, including job queues, audio decoders, hardware workers, and health monitoring.
OpenAI Whisper Reference Guide ; OpenAI Whisper Model Card ; Whisper Code and Weights License

Whisper Model Sizes and Hardware

Whisper offers six main model sizes, alongside dedicated .en versions for English that are often useful at tiny and base sizes. While multilingual non-turbo models support translation into English, turbo is not trained for translation. Earlier checkpoints like large-v1 and large-v2 remain available, so specify an explicit checkpoint. Because quality and speed vary across hardware and audio, always test directly; there is no guaranteed model ranking.

Inference VRAM Requirements by Model Size
ModelVariantsApprox. VRAMWhere to Start
tinytiny, tiny.en~1 GBTest when minimal memory is essential and lower accuracy on noisy audio is acceptable.
basebase, base.en~1 GBTest on modest hardware when memory is tight and tiny produces too many transcription errors.
smallsmall, small.en~2 GBTest when limited GPU memory allows more compute for everyday English or multilingual recordings.
mediummedium, medium.en~5 GBTest when mid-tier GPU memory is available for audio with mixed accents or multilingual speech.
largelarge-v3, large-v2, large-v1~10 GBTest when available VRAM supports larger compute budgets for difficult acoustics or speech translation.
turboturbo (multilingual only)~6 GBTest when prioritizing faster multilingual transcription on capable GPUs without needing speech translation into English.

Stated inference VRAM figures are approximate estimates from upstream documentation for running inference on a single GPU. They do not represent CPU RAM, audio conversion buffers, or total system memory. Whisper runs on CPUs or GPUs; measure throughput on your target CPU or GPU.

OpenAI Whisper Reference Guide

Choose How to Use Whisper

Choose the deployment model that best aligns with your team's privacy rules, technical resources, and budget.

1. Local Development and Testing

Install the reference package locally using Python, PyTorch, and ffmpeg. The MIT license incurs no software fees, but hardware, compute time, and setup labor remain real costs. This approach enables private testing on your workstation before allocating dedicated servers.

Read the Official Setup Guide

2. Self-Hosted Private Production

Choose CPU or GPU workers for control over data location, offline operation, or customization. Pre-download weights and dependencies, verify surrounding integrations stay offline, and manage your own queues, scaling, and ongoing maintenance.

Compare Self-Hosting Costs

3. Managed Hosted Whisper APIs

Use a managed cloud API to bypass server maintenance and hardware scaling. Verify the provider's specific model checkpoint, upload file limits, required output formats, data retention policies, and billing terms. Do not assume hosted whisper-1 equals large-v3 or turbo. Client integration, retry logic, and accuracy evaluation remain your responsibility.

Compare OpenAI API Alternatives

4. Speech is Cheap for Recorded Audio

Speech is Cheap is a managed API for recorded audio. Using it does not require installing Whisper or selecting its checkpoint. It provides JSON, SRT, and VTT outputs with optional word timestamps and speaker labels. Check supported languages and test your audio first; client integration is still needed.

View Pricing and Plans
OpenAI Whisper Reference Guide ; OpenAI Hosted Whisper Model ; Speech is Cheap Job Creation

How Whisper Processes Audio

Audio Ingestion and Windows

Whisper relies on ffmpeg to decode various audio formats into 16 kHz mono waveforms. Because the reference decoder downmixes audio into a single channel, preserve separate speaker channels when needed before transcription. Whisper processes audio in 30-second sliding windows, which represent internal neural network chunk sizes rather than audio length limits. Local software can process long audio files sequentially, constrained by system memory, decoding speed, and compute capacity.

Transcription and Outputs

Whisper can transcribe speech and identify language; multilingual non-turbo models translate speech into English. In Python, transcribe returns text, segments, and language, with optional word timestamps. The CLI exports TXT, VTT, SRT, TSV, and JSON formats.

Production Infrastructure

Whisper alone is not a turnkey production service or live streaming platform, and it lacks native speaker diarization. Operating at scale requires job queues to handle spikes, automated retries for transient failures, and system monitoring across worker instances. Hardware provisioning, maintenance costs, and architecture trade-offs are detailed in the resources linked below.

Self-Hosted Whisper vs API

Whisper File Processing ; OpenAI Whisper Model Card

Technical Limitations to Consider

Invented Text and Repetition Loops

Whisper can occasionally generate invented words or repeat phrases in loops during periods of silence, background noise, or degraded audio. Review transcripts when errors could affect a decision.

Variable Multilingual and Accent Accuracy

Accuracy varies across languages, accents, recording devices, and acoustic environments. While languages with abundant training data often yield fewer errors, high-resource languages are not universally better in every recording condition. Larger models do not eliminate transcription errors across unfamiliar terminology or non-standard accents.

Speakers and Live Audio

Whisper transcribes speech without separating speakers or assigning speaker labels, requiring a separate diarization tool if speaker attribution is needed. It is also designed for recorded audio files rather than real-time live audio, so low-latency streaming requires custom buffering and streaming architectures.

OpenAI Whisper Reference Guide ; OpenAI Whisper Model Card

Open-Source Status and Maintenance

As verified on 2026-09-17, the official OpenAI Whisper repository had its main branch commit on 2026-08-31, and its latest tagged release was v20250625, published 2025-06-26. This source snapshot carries no promised support schedule or roadmap for future updates. Teams deploying Whisper handle their own upgrade testing, dependency pinning, and ongoing maintenance.

Whisper v20250625 Release ; OpenAI Whisper Reference Guide

Frequently Asked Questions

Is Whisper Free to Use?

The open-source code and model weights are licensed under the MIT license with no software fees. In practice, running Whisper requires hardware, power, and maintenance time, while commercial cloud APIs bill based on usage.

Can Whisper Transcribe Live Audio in Real Time?

Whisper processes audio in 30-second windows and is designed for recorded audio files rather than turnkey streaming. While developers can construct custom streaming pipelines or buffer audio around the model, the reference implementation does not provide an out-of-the-box, low-latency live streaming solution.

Does Whisper Provide Speaker Diarization?

No. Whisper does not separate speakers or assign speaker labels. It generates transcript text and timestamps. Adding speaker labels requires pairing Whisper with a separate speaker diarization tool.

Can Whisper Run Offline Without Internet Connectivity?

Whisper executes locally without internet access once the software, dependencies, and model weights are downloaded. However, offline operation depends on the full surrounding workflow, including verifying that local audio pipelines and adjacent application services do not transmit data over external networks.

Is the Largest Model Always the Best Choice?

Selecting a model size is a trade-off to test on your own audio, languages, and hardware. The ~10 GB VRAM listed for large-v3 reflects GPU memory rather than CPU RAM or total machine memory. Whisper runs on either CPUs or GPUs; measure throughput on your target CPU or GPU. Smaller models or turbo may provide acceptable accuracy with lower compute requirements.

Use a Managed Transcription API

If you prefer not to manage local hardware or server infrastructure, evaluate a managed API using your own recorded audio to compare accuracy, turnaround time, and total cost.

Source Material and Verification

All specifications reference the official OpenAI Whisper repository, verified on 2026-09-17. External provider terms and pricing vary independently, and no independent runtime benchmarks were conducted for this guide.