Speech-to-Text & TTS Cost Calculator — Compare API Pricing

Quick Start Guide

  1. Select a service & model: Search the model catalog — Amazon, AssemblyAI, Cartesia, Deepgram, ElevenLabs, Google Cloud, Microsoft Azure, OpenAI, Rev AI, and Speechmatics — grouped by provider, with STT and TTS models both listed.
  2. Switch to Compare when needed: Benchmark the same workload across multiple providers of the same type (STT or TTS) side by side.
  3. Optionally load a scenario preset: Prefill a typical workload — weekly podcast transcription, daily support calls, an audiobook voice batch, or voice assistant responses — or leave it as Custom to enter your own values.
  4. Set your workload: STT uses audio duration in minutes; TTS uses character count.
  5. Tune run assumptions: Adjust runs per month, billing increment, and optional surcharges for realtime processing, speaker diarization, or content moderation.
  6. Review the cost: See base cost, surcharge impact, and the projected monthly total.

Understanding Speech API Pricing

Per-minute vs. per-character billing

Speech-to-text providers bill by audio duration processed — typically per minute. Text-to-speech providers bill by the number of characters (or words) you send to be synthesized, usually per 1,000 characters. These are fundamentally different pricing models, so switching from an STT to a TTS model below changes both the input field and how the rate is expressed.

Billing increments round up your usage

Most providers round usage up to a fixed increment before billing — for example, a 15-second minimum per STT request, or per-1,000-character blocks for TTS. A 3-second clip billed in 15-second increments still costs the same as a full 15 seconds. This calculator applies that rounding, so the billed usage shown can be higher than what you typed in.

Why realtime streaming costs more than batch

Processing audio as it arrives (realtime/streaming transcription for live captions or voice assistants) requires dedicated compute held open for the duration of the session, unlike batch processing a pre-recorded file after the fact. Providers typically charge a premium — often 20-80% above the batch rate — for this. Use the Realtime Premium toggle in Run Assumptions to approximate it.

Stacking the optional surcharges

Realtime streaming, speaker diarization and profanity filtering are each priced as a percentage on top of the base rate, and they ADD rather than compound. A 60-minute Whisper job costs $0.36 at the base rate; turn on realtime at 25% and it is $0.45, diarization alone at 12% makes it $0.4032, and both together apply 37% for $0.4932 — not the $0.504 you would get by multiplying the two surcharges. The distinction is small on one job and not small across a month of them, which is what the monthly-runs figure is for.

Standard vs. neural voice quality

For text-to-speech, "Standard" voices use older, less compute-intensive synthesis and cost less. "Neural" voices sound noticeably more natural but run on more expensive models — commonly 3-4x the per-character rate of Standard. Toggle Voice Quality above to compare the two where a provider offers both.

Features

10-Provider Model Catalog: Amazon, AssemblyAI, Cartesia, Deepgram, ElevenLabs, Google Cloud, Microsoft Azure, OpenAI, Rev AI, and Speechmatics, spanning both STT and TTS models.

STT & TTS Support: Calculate costs for both speech-to-text and text-to-speech services, switching input fields automatically based on the selected model.

Compare Mode: Benchmark identical workload assumptions across multiple providers of the same type at once.

Scenario Presets: Four ready-made workload templates — weekly podcast transcription, daily support calls, an audiobook voice batch, and voice assistant responses.

Voice Quality Tiers: For supported TTS models, compare Standard vs. Neural voice pricing.

Run Assumptions: Billing increment, monthly run cadence, and optional surcharge toggles for realtime premium, speaker diarization, and content moderation.

Custom Pricing: Override the catalog rate with your own negotiated contract price.

Export Data: Export your speech cost calculations to JSON, CSV, or PDF format for project planning and budgeting.

Common Use Cases

Weekly Podcast Transcription: 60-minute episodes with speaker diarization, transcribed 4 times a month — a typical STT workload for a recurring podcast.

Daily Support Call Review: 45 minutes of realtime STT per business day, 22 days a month — sized for ongoing call-quality or compliance review.

Audiobook Voice Batch: roughly 120,000 characters of TTS narration per month — a manuscript-scale voice generation run.

Voice Assistant Responses: 25,000 characters per run across 30 runs a month with realtime premium and neural voice quality — a production voice-interface workload.

Frequently Asked Questions

OpenAI Whisper (speech-to-text) is priced at $0.006 per minute of audio processed, rounded to the nearest second. For example, 100 minutes of transcription costs $0.60. Enter your expected monthly audio volume to project total monthly spend.
TTS-1 is optimized for low latency (real-time applications), while TTS-1-HD offers higher audio quality but may be slower. HD costs twice as much per character.
Pricing varies by provider and usage tier across the ten providers in this calculator — Amazon, AssemblyAI, Cartesia, Deepgram, ElevenLabs, Google Cloud, Microsoft Azure, OpenAI, Rev AI, and Speechmatics. At low volumes, Whisper and Google Speech-to-Text are often competitive. At high volumes, batch discounts and custom contracts can significantly reduce costs. Use Compare mode to run the same monthly audio volume across providers and see ranked costs.
Enter the average audio duration (for STT) or character count (for TTS) per request and the number of requests per month. The calculator multiplies the per-unit cost by your monthly volume and projects the total spend. Switch to Compare mode to see how the estimate changes across different providers.
No — they add. Realtime at 25% and diarization at 12% apply as a combined 37% on the base rate, not 1.25 × 1.12. On a $0.36 job that is $0.4932 rather than $0.504. The same holds for the profanity-filter surcharge, so switching several on tells you the total percentage rather than making you work out the interaction.
Most providers round usage up to a fixed billing increment before charging — for example, a 15-second minimum per STT request, or per-1,000-character blocks for TTS. A short clip or a small batch of text can still be billed at the full increment. Run Assumptions shows and lets you adjust this rounding.

Last updated