SunsettingAI Audio

whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-transcribe-diarize

by OpenAI

Announced

2026-08-26

Complete

2027-02-26

Refund window closes

—

Refund status

No refund applicable

OpenAI's four hosted transcription models lose API access on February 26, 2027. The deprecations page records the announcement on August 26, 2026 for `whisper-1`, `gpt-4o-transcribe`, `gpt-4o-mini-transcribe` and `gpt-4o-transcribe-diarize`, and names `gpt-transcribe` or `gpt-live-transcribe` as the replacement for all four. That is six months of notice, which matches OpenAI's stated minimum for generally available models. The part that matters if you caption video: as of September 28, 2026 OpenAI's own transcription guide still routes word timestamps, SRT and VTT subtitle output, and English translation of a finished recording to `whisper-1`, and speaker-labeled transcripts to `gpt-4o-transcribe-diarize`. Those are four documented capabilities pointing at models with a shutdown date, and the guide does not yet name a post-sunset path for them.

Refund flow

  1. 1

    Not applicable. These are pay-per-minute audio models with no pre-paid balance tied to a model string, so there is no refund category to open.

  2. 2

    If you are billed for one of these four model strings for usage dated after February 26, 2027, treat it as a billing error rather than a refund request: open a ticket at help.openai.com with the invoice line and the request IDs.

  3. 3

    The thing to check before the date is not a refund, it is whether your caption pipeline still has a documented model. Transcription is priced per minute of audio, so a switch changes unit cost as well as output format.

Migration path

Recommended successor

gpt-transcribe for files, gpt-live-transcribe for live audio

Why this is the closest fit

These are the replacements OpenAI names itself, on every row of the August 26, 2026 deprecation table, and the two models its current transcription guide recommends for new integrations.

What differs from the original

The split is new. Where `whisper-1` was one endpoint for everything, the replacement path asks you to choose first: `gpt-transcribe` for a completed file or a bounded request, `gpt-live-transcribe` for audio arriving from a microphone or a call. Streaming the output of a finished file no longer requires a Realtime session. The gap is output format rather than accuracy: the guide documents word timestamps and `srt` and `vtt` subtitle generation against `whisper-1`, and the English-translation endpoint against `whisper-1`, so a pipeline that renders subtitle files from OpenAI transcripts has no drop-in documented successor today. Speaker diarization has the same problem one step removed, since the guide sends it to `gpt-4o-transcribe-diarize`, which is on the same February 26, 2027 list.

Alternatives by use case

You call whisper-1 for plain file transcription

gpt-transcribe, the model OpenAI recommends for completed recordings. Re-test on your own audio before switching, because per-minute price and error profile both change

You generate SRT or VTT subtitle files for video

Do not assume a string swap. Verify on the current speech-to-text guide whether gpt-transcribe has gained subtitle output before February 26, 2027, and keep a self-hosted Whisper build as the fallback, since the open weights do not expire with the hosted endpoint

You need speaker labels

gpt-4o-transcribe-diarize still works today but shuts down on the same date, so treat it as a temporary path and re-check the guide rather than migrating onto it

You transcribe live audio from a call or microphone

gpt-live-transcribe, which is the recommended Realtime path and the cleaner half of this migration

You translate finished recordings into English

The audio translations endpoint is documented against whisper-1 only, so plan on either a two-step transcribe-then-translate flow or a self-hosted Whisper for this one

What this tool meant

Whisper is the model that made speech-to-text a commodity. When OpenAI open-sourced the weights in September 2022 and put a hosted endpoint behind `whisper-1` in March 2023, transcription went from a per-minute vendor service to something any developer could call for cents or run on a laptop. Most AI captioning, subtitle and meeting-notes tooling built between 2023 and 2025 has Whisper somewhere in it. What is retiring on February 26, 2027 is the hosted endpoint, not the model: the open weights stay downloadable and runnable, which makes this an unusually soft deprecation compared with a closed model like Sora 2, where the shutdown removed the only way to run it at all. The transferable lesson is narrower than "hosted models die." It is that capabilities documented against a single model are the fragile part of an integration. Accuracy migrates easily, because a newer model is usually better. Output formats do not, because a format is a feature someone has to reimplement, and subtitle files, word-level timestamps and speaker labels are exactly the features that get dropped when a provider redesigns an API around conversation rather than around media. If your product renders text onto video, read the capability table on the provider's guide, not just the model list on its pricing page.

Sources

Spent credits on AI video that failed?

AVA automates the refund flow across every video provider

Free Chrome extension. Detects the failure mode, captures evidence, drafts the refund email with the technical term and Generation ID. Click send.

Other shutdowns in AI Audio