tts-1, tts-1-hd, gpt-4o-mini-tts
by OpenAI
Announced
2026-10-01
Complete
2027-01-06
Refund window closes
—
Refund status
No refund applicable
On October 1, 2026 OpenAI deprecated its text-to-speech models: `tts-1`, `tts-1-hd`, `gpt-4o-mini-tts-2025-03-20` and `gpt-4o-mini-tts-2025-12-15` are removed from the API on January 6, 2027. The deprecations page names one replacement for all four, `gpt-realtime-2.1-mini`, and points to the Realtime API guide to plan the migration. That is just over three months of notice, which matches the "at least three months" the announcement states. The detail that matters if you generate voiceover files for video: as of October 3, 2026, every model being retired lists Speech generation (`v1/audio/speech`) as a supported endpoint, and the `gpt-realtime-2.1-mini` model page lists Realtime (`v1/realtime`) as its only supported endpoint. The replacement is reached through a different API, not by changing the model string in the call you already make.
Refund flow
- 1
Not applicable. These are pay-as-you-go API models with no pre-paid balance tied to a model string, so there is no refund category to open.
- 2
If you are billed for tts-1, tts-1-hd or gpt-4o-mini-tts usage dated after January 6, 2027, treat it as a billing error rather than a refund request: open a ticket at help.openai.com with the invoice line and the request IDs.
- 3
Check the price units before you budget the move, because they are not the same. As listed on OpenAI's model pages on October 3, 2026: tts-1 is $15 per 1M characters and tts-1-hd is $30 per 1M characters. gpt-4o-mini-tts is $0.60 per 1M text input tokens and $12 per 1M audio output tokens. gpt-realtime-2.1-mini is $0.60 per 1M text input tokens and $20 per 1M audio output tokens. For a gpt-4o-mini-tts workload the listed audio output rate goes from $12 to $20 per 1M tokens. For tts-1 and tts-1-hd, which bill per character, there is no direct comparison on the pages themselves; meter a real narration job on the new model to get your number.
Migration path
Recommended successor
gpt-realtime-2.1-mini (Realtime API)
Why this is the closest fit
It is the substitute OpenAI names on every row of the October 1, 2026 text-to-speech deprecation, and its model page lists audio as an output modality, so it can produce spoken audio from text.
What differs from the original
The call shape changes. The retiring models are used with the Speech endpoint: one POST with model, voice and input text, and the response body is the audio file (mp3 by default, with opus, aac, flac, wav and pcm also documented). gpt-realtime-2.1-mini does not list that endpoint; it lists only the Realtime API, which OpenAI documents as a session you connect to over WebRTC, WebSockets or SIP. A pipeline that renders a voiceover file per video therefore needs new session handling and audio capture code, not a find-and-replace. Voices can change too: the text-to-speech guide says tts-1 and tts-1-hd support alloy, ash, coral, echo, fable, onyx, nova, sage and shimmer, and notes that the Realtime API's voice set is slightly different. Finally, as of October 3, 2026 the text-to-speech guide's own code samples still use gpt-4o-mini-tts, one of the models being retired, so do not take the guide's examples as the post-January path.
Alternatives by use case
You render one voiceover file per video in a batch job
Prototype a Realtime API session on gpt-realtime-2.1-mini now and confirm you can capture a complete audio file from it at the quality you ship, since the Speech endpoint call you use today has no listed equivalent on the named replacement
You use gpt-4o-mini-tts for its tone instructions ("speak in a cheerful tone")
Re-test your instruction prompts on gpt-realtime-2.1-mini before January 6, 2027. Delivery style is set differently in a realtime session, and output will not be identical
You chose tts-1 for low cost on long narration
Run one full-length script through the new model and read the audio output token count on the invoice, because per-character and per-token prices cannot be compared on paper
Your brand voice depends on a specific tts-1 voice name
Check the voice exists in the Realtime API voice list before you migrate. If it does not, budget time to pick and approve a new voice
What this tool meant
tts-1 was the simplest way to put a voice on an AI-generated video: send text, get back an mp3. That simplicity is the thing this deprecation removes. OpenAI is folding text-to-speech into its realtime voice stack, which is built for live conversation rather than file rendering. For AI video pipelines the lesson is the same one the whisper-1 transcription deprecation taught: a named replacement is a replacement for capability, not for the endpoint, the price unit or the voice list you built around. Read the replacement's model page, not just the deprecation row, while there are still months left to do the work.
Sources
- OpenAI API deprecations page (2026-10-01 text-to-speech deprecation, January 6, 2027 shutdown, named replacement)
- OpenAI model page: GPT-Realtime-2.1 Mini (pricing, supported endpoints)
- OpenAI model page: TTS-1 (pricing, Speech endpoint, deprecated snapshot)
- OpenAI model page: TTS-1 HD (pricing)
- OpenAI model page: GPT-4o mini TTS (pricing, deprecated snapshots)
- OpenAI text-to-speech guide (Speech endpoint, output formats, voice lists)
- AVA record for the whisper-1 transcription shutdown (February 26, 2027)
Spent credits on AI video that failed?
AVA automates the refund flow across every video provider
Free Chrome extension. Detects the failure mode, captures evidence, drafts the refund email with the technical term and Generation ID. Click send.