Automatic speech recognition converts spoken audio into written text. Whisper performs three core tasks: transcribing speech in its original language, translating non-English speech directly into English text, and identifying the spoken language in a recording. While trained on diverse audio, transcription accuracy varies across speakers, background noise, recording equipment, and technical terminology.
Evaluating Whisper requires distinguishing the trained neural network from the software used to run it. Developers separate model weights from local runtime libraries, self-hosted servers, and commercial cloud APIs. Choosing an approach depends on data privacy needs, technical resources, and budget.
- Model Weights and Checkpoints
- Trained neural network parameters released under the MIT license across multiple sizes, converting speech patterns into written text.
- Reference Software Implementation
- The official OpenAI Python package providing command-line tools and PyTorch code to run inference on CPU or GPU hardware.
- Self-Hosted Production Stack
- The surrounding software required to run Whisper reliably at scale, including job queues, audio decoders, hardware workers, and health monitoring.