How Are Costs Calculated?
Default pipeline throughput of 1,200 audio minutes per active GPU hour (20x real-time) serves as a baseline planning assumption shared across all models, not an empirical benchmark of specific model or runtime performance.
Full pipeline throughput encompasses end-to-end processing, including audio decoding, CPU normalization, optional word timestamps, speaker diarization, and retry overhead. Certain models natively include word timestamps within primary decoding. All baseline cost figures are illustrative starting estimates and fully editable, rather than formal vendor quotes.
Cloud elastic hosting models compute expense by dividing monthly audio minutes by the product of pipeline throughput and GPU utilization, represented as M / (T * u), and multiplying by the hourly GPU rate. Billed time accounts for idle capacity once within that utilization factor, assuming elastic provisioning scales to meet incoming demand.
Owned hardware costs are calculated as N * (purchase / amortMonths + power) for N dedicated GPUs. Monthly processing capacity is modeled as 730 * N * u * T based on a standard 730-hour operating month. Standby capacity and reserve margins must be budgeted through utilization targets or additional hardware, as owned clusters offer no automatic redundancy.
Self-hosted infrastructure adds fixed monthly overhead for CPU ingestion workers, job queues, and observability, plus storage and egress calculated as M * storageRate. Shared application expenses outside speech decoding are equally excluded from both options. Engineering labor applies a single hourly rate to both approaches: self-hosted labor is laborRate * (setupHours / setupMonths + opsHours), while API labor is laborRate * (apiSetupHours / setupMonths + apiOpsHours). Setting labor hours to zero models incremental capacity on an existing team, not free engineering.
Speech is Cheap API expenses reflect the lowest available rate between pay-as-you-go and monthly subscription tiers from maintained pricing data, applying optional add-on fees across all processed minutes. In examples, API billing rounds each audio file up to the next whole minute and excludes taxes, volume credits, or custom negotiations. Projected break-even identifies the first positive whole minute where total self-hosted costs undercut API billing inclusive of engineering effort on both sides, bounded by 100 million minutes or owned hardware capacity; achieving break-even does not guarantee self-hosting remains cheaper at higher volumes.
How Is Capacity Calculated?
PAYGFees and subscriptionFees are the full plan bills including all selected add-ons; choose the lower amount rather than a price per minute. All other formulas use GPU hours and monthly minutes.
- Cloud Compute Cost
- cloudCost = (M / (T * u)) * gpuRate
- Owned Hardware Cost
- ownedCost = N * ((purchase / amortMonths) + power)
- Owned Hardware Capacity
- capacity = 730 * N * u * T
- Self-Hosted Engineering Labor
- selfLabor = laborRate * ((setupHours / setupMonths) + opsHours)
- API Integration Labor
- apiLabor = laborRate * ((apiSetupHours / setupMonths) + apiOpsHours)
- Full Pipeline Totals
- totalSelfHosted = computeCost + fixed + (M * storageRate) + selfLabor; totalAPI = min(PAYGFees, subscriptionFees) + apiLabor
- M
- Monthly audio minutes processed
- T
- Pipeline throughput in audio minutes per active GPU hour
- u
- Target GPU capacity utilization fraction (0 to 1.0)
- N
- Number of owned GPUs in cluster
- gpuRate
- Hourly cloud GPU rental rate in USD
- purchase
- Hardware purchase price per GPU in USD
- amortMonths
- Hardware depreciation amortization schedule in months
- power
- Monthly hosting, power, and cooling cost per GPU in USD
- capacity
- Maximum monthly audio processing capacity in minutes
- fixed
- Monthly fixed infrastructure cost for ingestion workers, queues, and monitoring in USD
- storageRate
- Storage and egress fee per audio minute in USD
- laborRate
- Blended engineering hourly rate in USD
- setupHours
- Initial engineering setup hours for self-hosting
- setupMonths
- Setup labor amortization period in months
- opsHours
- Ongoing monthly engineering maintenance hours for self-hosting
- apiSetupHours
- Initial engineering setup hours for API integration
- apiOpsHours
- Ongoing monthly engineering maintenance hours for API integration
- computeCost
- Monthly GPU compute expense for cloud elastic or owned hardware in USD
- totalSelfHosted
- Total monthly self-hosted pipeline cost in USD
- totalAPI
- Total monthly managed API cost in USD
Speech is Cheap Pricing ; Speech is Cheap Job Creation