$5 free credits when you sign up Claim now
Marlin 2b now available Test it!
MiniMax H3 now available! Test it!
Whisper API Pricing per Minute: deAPI vs OpenAI
admin Oct 2, 2026 7 min read

Whisper API Pricing per Minute: deAPI vs OpenAI

Every hour of audio your users upload lands on your transcription bill. If you run a podcast tool or a meeting recorder, that bill grows with your user count, and the per-minute rate decides whether the feature pays for itself.

This guide compares Whisper API pricing on OpenAI and deAPI, with prices checked on September 30, 2026. It covers cost per minute, per file, and per month, plus the code to point your OpenAI SDK at deAPI, a cheap transcription API running open-source Whisper Large V3.

ModelProviderPrice1 hour of audio
Whisper Large V3 / V3 CT2 (files, YouTube)deAPI$0.00078/min + $0.005 per job$0.052
Whisper Large V3 CT2 with speaker labelsdeAPIbase price +50%$0.078
gpt-4o-mini-transcribeOpenAI$0.003/min$0.18
gpt-transcribeOpenAI$0.0045/min$0.27
whisper-1OpenAI$0.006/min$0.36
gpt-4o-transcribe-diarizeOpenAI$0.006/min$0.36

Whisper API cost per minute on each side

OpenAI’s transcription API pricing is a flat rate per minute of audio. A 10-second clip on whisper-1 costs a tenth of a cent, a one-hour file costs $0.36.

deAPI splits the price in two: $0.046875 per hour of audio (about $0.00078 per minute) plus a flat $0.005 per job. The flat fee makes up most of the bill on short clips and almost nothing on long recordings.

Past about 2.3 minutes of audio, deAPI costs less than both OpenAI models, and the gap widens with every minute. On shorter clips the job fee dominates, and the batching approach in the next section takes care of that.

File lengthdeAPIwhisper-1gpt-4o-mini-transcribe
30 seconds$0.0054$0.0030$0.0015
5 minutes$0.0089$0.0300$0.0150
15 minutes$0.0167$0.0900$0.0450
1 hour$0.0519$0.3600$0.1800
3 hours$0.1456$1.0800$0.5400

At one hour, deAPI costs about 7 times less than whisper-1 and 3.5 times less than gpt-4o-mini-transcribe, which makes it the cheapest transcription API in this comparison.

Monthly transcription cost for four workloads

WorkloadAudio per monthdeAPIwhisper-1gpt-4o-mini-transcribe
Weekly podcast, 4 × 1 h4 h$0.21$1.44$0.72
Team meetings, 40 × 30 min20 h$1.14$7.20$3.60
YouTube archive, 100 × 15 min25 h$1.67$9.00$4.50
Voice notes, 1,000 × 45 s, batched daily12.5 h$0.74$4.50$2.25

The voice-notes row assumes batching. Sent one by one, a thousand 45-second clips pay the $0.005 job fee a thousand times and the month comes to $5.59. Stitch each day’s clips into one file with ffmpeg, transcribe it once, and split the transcript back by segment timestamps. A day of notes runs about 25 minutes, which fits the 20 MB upload limit as a 64 kbps MP3, and 30 jobs a month cost $0.74.

Speaker labels cost less too

Speaker diarization answers “who said what” in meetings and multi-host podcasts. On deAPI it runs on Whisper Large V3 CT2 and adds 50% to the base price, which puts an hour of labeled audio at $0.078. OpenAI’s diarization model, gpt-4o-transcribe-diarize, costs $0.36 per hour, about 4.6 times more.

Word-level timestamps carry the same 50% surcharge on deAPI. Segment timestamps are free on both Whisper models. For output formats and a full diarization walkthrough, see Whisper Large V3 CT2: Transcription That Knows Who Said What.

A drop-in Whisper API alternative for the OpenAI SDK

deAPI exposes an OpenAI-compatible endpoint, so existing code needs a new base URL, a deAPI key, and a model name. There are no aliases for whisper-1 or gpt-4o-transcribe, so you pass WhisperLargeV3 or WhisperLargeV3Ct2 directly.

from openai import OpenAI

client = OpenAI(
    api_key="dpn-sk-...",                 # deAPI key from app.deapi.ai
    base_url="<https://oai.deapi.ai/v1>",
    timeout=900.0,                        # long files can take several minutes
)

with open("recording.mp3", "rb") as f:
    transcript = client.audio.transcriptions.create(
        model="WhisperLargeV3Ct2",
        file=f,
        response_format="verbose_json",
    )

print(transcript.text)
for seg in transcript.segments:
    print(f"[{seg.start:.1f}-{seg.end:.1f}] {seg.text}")

We ran this on a 10-minute public-domain LibriVox recording (4.9 MB MP3). Whisper Large V3 CT2 returned 23 timestamped segments in 10 seconds and charged $0.01296, exactly what the price endpoint quoted before the run. The same file on Whisper Large V3 took 16 seconds.

The 900-second timeout matters because the OpenAI SDK gives up at 600 seconds by default, and the deAPI gateway holds the connection open until the job finishes. More on the compatibility layer in Use Your OpenAI SDK with deAPI.

Transcribe straight from a URL

OpenAI’s transcription endpoint accepts file uploads up to 25 MB, which means downloading and compressing the media yourself first. deAPI’s native endpoint also takes a public link from YouTube, X, Twitch, Kick, TikTok, or X Spaces and fetches the audio on its side.

import os
import requests

resp = requests.post(
    "https://api.deapi.ai/api/v2/audio/transcriptions",
    headers={"Authorization": f"Bearer {os.environ['DEAPI_API_KEY']}"},
    json={
        "source_url": "https://www.youtube.com/watch?v=eG3mQzYbwIY",
        "model": "WhisperLargeV3",
        "include_ts": True,
    },
)
print(resp.json())

The request returns a request_id. Poll GET <https://api.deapi.ai/api/v2/jobs/{request_id}> until status is done, then download the transcript from result_url.

We sent the first episode of NASA’s Houston We Have a Podcast, 46 minutes long. The job finished in 73 seconds of polling and cost $0.041. The transcript opens like this:

[0:00 – 0:02] Houston, we have a podcast. > [0:03 – 0:27] Welcome to the official podcast of the NASA Johnson Space Center, Episode 1, International Space Station. I’m Gary Jordan, and I’ll be your host today.

The rate depends on the platform. YouTube links cost the same as uploaded files. Twitch, Kick, and X links cost about $0.26 per hour of audio, five times the file rate and still below whisper-1. TikTok links work only with Whisper Large V3 CT2 and are priced per video, from $0.015 under a minute to $0.105 over 10 minutes. For a full YouTube walkthrough, see How to Transcribe YouTube Videos with AI.

Check the price before you run a job

The price endpoint quotes a job from its duration, a file, or a URL, and charges nothing:

curl -X POST <https://api.deapi.ai/api/v2/audio/transcriptions/price> \
  -H "Authorization: Bearer $DEAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"duration_seconds": 7200, "model": "WhisperLargeV3Ct2", "diarize": true}'

This returns {"data":{"price":0.148125}} for two hours with speaker labels. Send the same flags you plan to use on the real job, since diarize and word timestamps raise the price and a quote without them comes in low. Price checks run on a separate rate-limit bucket, so quoting every upload in your app does not eat into your transcription quota.

FAQ

What is the cheapest transcription API for long recordings?

Of the APIs compared here, deAPI: $0.052 for an hour-long file or YouTube video. The same hour costs $0.36 on OpenAI whisper-1 and $0.18 on gpt-4o-mini-transcribe.

How much is the OpenAI Whisper API per minute?

Whisper-1 costs $0.006 per minute, gpt-4o-mini-transcribe $0.003. On deAPI, Whisper Large V3 costs about $0.00078 per minute plus $0.005 per job.

Is there a free tier for the Whisper API on deAPI?

New accounts get a $5 bonus with no card required. At $0.052 per hour-long file, that covers about 96 hours of audio. Accounts registered with disposable email addresses do not receive the bonus.

Are there rate limits on the free tier?

Accounts without a payment can send 1 transcription request per minute and 10 per day. Any top-up switches the account to 300 requests per minute with no daily cap, and the remaining bonus carries over.

What happens if a job fails?

Failed jobs are normally refunded, and the job status shows refunded: true with a price of 0. Audio with no speech still completes and returns an empty transcript, so check for that case in your code.

Which Whisper model should I pick on deAPI?

Both cost the same at the base rate. Whisper Large V3 CT2 adds speaker labels, word-level timestamps, and TikTok support, so it is the better default. Plain Whisper Large V3 fits when you only need text with segment timestamps.


Get a price for your own files. Send one of your recordings to the price endpoint, then transcribe it on the $5 signup credit: create a deAPI account.

No subscription No credit card required

Start building with AI in under a minute

Access all models from this article through a single REST API. Start with $5 free credits — no subscription, no credit card.

Migration assistance available talk to an engineer