OpenAI: GPT-4o Mini Transcribe (Audio to Text)
openai/gpt-4o-mini-transcribe/audio-to-text
The gpt-4o-mini-transcribe model is a highly efficient speech-to-text solution designed to deliver accurate audio transcriptions while optimizing for speed and resource consumption. This model offers significant improvements in word error rate and language recognition, making it particularly effective in scenarios involving accents, noisy environments, and varying speech speeds. gpt-4o-mini-transcribe is ideal for applications that require quick and reliable transcription services. gpt-4o-mini-transcribe has been pretrained on specialized audio-centric datasets, which include diverse and high-quality audio samples, ensuring a deep understanding of speech nuances. This model supports a substantial context window of 16,000 tokens, allowing it to process longer audio inputs effectively. With a maximum output of 2,000 tokens, gpt-4o-mini-transcribe can generate detailed and comprehensive transcriptions. The training process incorporates rigorous enhancement techniques, including supervised fine-tuning and reinforcement learning, to optimize performance and accuracy.
Model specifications
- Input
- audio
- Output
- text
- Context
- 10,000 tokens
- Max output
- 10,000 tokens
- Input price