Speech-to-Text Infographics

What a Language List Cannot Tell You

A historical look at the original Whisper training mix, and why language coverage is only a starting point.

Speech is Cheap

Historical case study
Whisper, 2022

Multilingual Does Not Mean Equally Trained

The original Whisper training mix shows why a long language list is only the start of the story.

65%

English audio
English transcript

18%

Non-English audio
English translation

17%

Non-English audio
Transcript in its language

680k training hours

The model card reports that transcription performance in a language correlates with its training-data volume.

“Multilingual” is a capability label, not an equal-accuracy guarantee.

A historical lesson from the 2022 release. This graphic does not describe the Speech is Cheap backend.

Source: OpenAI Whisper model card, training-data and limitations sections, checked 2026-09-15. Shares and hours are approximate as reported. The source describes uneven language performance; correlation does not establish causation.

Evaluate the Language Your Users Speak

The original 2022 Whisper model card describes a training mix with far more English transcription than non-English transcription. It also reports a relationship between training-data volume and transcription performance by language. This historical example explains why a supported-language list cannot establish accuracy for your recordings. It does not describe the Speech is Cheap backend.

Put Your Recordings to Work

Review included minutes, overage and optional features. Choose the Speech is Cheap API plan that fits your recording workflow.