OpenAI: Whisper-1 (Audio to Text)
openai/whisper-1/audio-to-text
The Whisper models are trained for speech recognition and translation tasks, capable of transcribing speech audio into the text in the language it is spoken (automatic speech recognition) as well as translated into English (speech translation). Researchers at OpenAI developed the models to study the robustness of speech processing systems trained under large-scale weak supervision. The model version 001 corresponds to whisper large v2. Max request data size: 25mb of audio can be converted from speech to text per API request.
Model specifications
- Input
- audio
- Output
- text
- Context
- 10,000 tokens
- Max output
- 10,000 tokens
- Input price
- $0.006 / 1M tokens
- Output price
- $0 / 1M tokens
- Released
- 2026-02-18
Capabilities
- Streaming
Provider pricing, discounts and data privacy
Compare effective provider prices, published discounts, regions, retention policies, training use, compliance, and privacy links by service tier.
Standard service tier
1 available provider · tier input average $0.006 / 1M tokens
Azure
Tier: Standard · Region: US
Pricing
- Input
- $0.006 / 1M tokens
No provider discount is currently published.
Data privacy and compliance
- Region
- US
- Zero data retention
- Yes
- Data retention
- Zero retention
- Used for training
- No
- Data collection
- Moderated
- No
- GDPR compliant
- Yes
- HIPAA compliant
- Yes
- SOC 2 certified
- Yes
- BYOK supported
- No
Privacy policy · Terms · Official website · Documentation · Status
Frequently asked questions
- What is OpenAI: Whisper-1 (Audio to Text)?
- The Whisper models are trained for speech recognition and translation tasks, capable of transcribing speech audio into the text in the language it is spoken (automatic speech recognition) as well as translated into English (speech translation). Researchers at OpenAI developed the models to study the robustness of speech processing systems trained under large-scale weak supervision. The model version 001 corresponds to whisper large v2. Max request data size: 25mb of audio can be converted from speech to text per API request.
- How much does OpenAI: Whisper-1 (Audio to Text) cost?
- Input costs start at $0.006 / 1M tokens and output costs start at $0 / 1M tokens. Provider-level prices vary by service tier.
- What is the context length of OpenAI: Whisper-1 (Audio to Text)?
- OpenAI: Whisper-1 (Audio to Text) supports a 10,000 token context window and up to 10,000 output tokens.
- What capabilities does OpenAI: Whisper-1 (Audio to Text) support?
- OpenAI: Whisper-1 (Audio to Text) supports Streaming.
- Which providers offer OpenAI: Whisper-1 (Audio to Text)?
- OpenAI: Whisper-1 (Audio to Text) is available from Azure.
- How do providers handle data privacy for OpenAI: Whisper-1 (Audio to Text)?
- 1 of 1 providers report that customer data is not used for training, and 1 offer zero-data-retention routing. Retention, compliance, and privacy-policy links are listed per provider.