OpenAI: GPT-4o Mini Transcribe (Audio to Text)

openai/gpt-4o-mini-transcribe/audio-to-text

The gpt-4o-mini-transcribe model is a highly efficient speech-to-text solution designed to deliver accurate audio transcriptions while optimizing for speed and resource consumption. This model offers significant improvements in word error rate and language recognition, making it particularly effective in scenarios involving accents, noisy environments, and varying speech speeds. gpt-4o-mini-transcribe is ideal for applications that require quick and reliable transcription services. gpt-4o-mini-transcribe has been pretrained on specialized audio-centric datasets, which include diverse and high-quality audio samples, ensuring a deep understanding of speech nuances. This model supports a substantial context window of 16,000 tokens, allowing it to process longer audio inputs effectively. With a maximum output of 2,000 tokens, gpt-4o-mini-transcribe can generate detailed and comprehensive transcriptions. The training process incorporates rigorous enhancement techniques, including supervised fine-tuning and reinforcement learning, to optimize performance and accuracy.

Model specifications

Input
audio
Output
text
Context
10,000 tokens
Max output
10,000 tokens
Input price
.25 / 1M tokens
Output price
$5 / 1M tokens
Released
2026-02-11

Provider pricing, discounts and data privacy

Compare effective provider prices, published discounts, regions, retention policies, training use, compliance, and privacy links by service tier.

Standard service tier

1 available provider · tier input average .25 / 1M tokens · tier output average $5 / 1M tokens

Azure

Tier: Standard · Region: US

Pricing
Input
.25 / 1M tokens
Output
$5 / 1M tokens

No provider discount is currently published.

Data privacy and compliance
Region
US
Zero data retention
Yes
Data retention
Zero retention
Used for training
No
Data collection
Moderated
Yes
GDPR compliant
Yes
HIPAA compliant
Yes
SOC 2 certified
Yes
BYOK supported
No

Frequently asked questions

What is OpenAI: GPT-4o Mini Transcribe (Audio to Text)?
The gpt-4o-mini-transcribe model is a highly efficient speech-to-text solution designed to deliver accurate audio transcriptions while optimizing for speed and resource consumption. This model offers significant improvements in word error rate and language recognition, making it particularly effective in scenarios involving accents, noisy environments, and varying speech speeds. gpt-4o-mini-transcribe is ideal for applications that require quick and reliable transcription services. gpt-4o-mini-transcribe has been pretrained on specialized audio-centric datasets, which include diverse and high-quality audio samples, ensuring a deep understanding of speech nuances. This model supports a substantial context window of 16,000 tokens, allowing it to process longer audio inputs effectively. With a maximum output of 2,000 tokens, gpt-4o-mini-transcribe can generate detailed and comprehensive transcriptions. The training process incorporates rigorous enhancement techniques, including supervised fine-tuning and reinforcement learning, to optimize performance and accuracy.
How much does OpenAI: GPT-4o Mini Transcribe (Audio to Text) cost?
Input costs start at .25 / 1M tokens and output costs start at $5 / 1M tokens. Provider-level prices vary by service tier.
What is the context length of OpenAI: GPT-4o Mini Transcribe (Audio to Text)?
OpenAI: GPT-4o Mini Transcribe (Audio to Text) supports a 10,000 token context window and up to 10,000 output tokens.
Which providers offer OpenAI: GPT-4o Mini Transcribe (Audio to Text)?
OpenAI: GPT-4o Mini Transcribe (Audio to Text) is available from Azure.
How do providers handle data privacy for OpenAI: GPT-4o Mini Transcribe (Audio to Text)?
1 of 1 providers report that customer data is not used for training, and 1 offer zero-data-retention routing. Retention, compliance, and privacy-policy links are listed per provider.

Browse all AI models