OpenAI: GPT-4o Mini TTS (Text to Audio) API Reference

About

GPT-4o mini TTS is a text-to-speech model built on GPT-4o mini, a fast and powerful language model.

Use it to convert text to natural sounding spoken text.

1. Calling the API

Setup your API Key

To begin using Infron, you first need to create an account and generate your API Key.

Once you have your API key, set it as an environment variable in your runtime environment. This allows your applications to securely access Infron’s API.

export INFRON_KEY="YOUR_API_KEY"

Submit a request

For long-running requests, you can poll for results.

curl https://media.onerouter.pro/v1/audios/generations \
  -H 'Authorization: Bearer YOUR_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "openai/gpt-4o-mini-tts/text-to-audio",
    "input": "Today is a wonderful day to build with Infron.",
    "voice": "alloy",
    "instructions": "esse incididunt",
    "response_format": "mp3",
    "speed": 1,
    "stream_format": "audio"
  }'

Response with queue task id

The response will look like this:

{
  "code": 200,
  "message": "success",
  "data": {
    "task_id": "f6c274b21ca3472c987e5f6383f3fa7b",
    "object": "audio",
    "model": "openai/gpt-4o-mini-tts/text-to-audio",
    "status": "created",
    "urls": {
      "query": "https://media.onerouter.pro/v1/audios/tasks/f6c274b21ca3472c987e5f6383f3fa7b"
    },
    "created_at": "2026-05-31 00:30:30.31"
  }
}

Fetch request status

You can fetch the status of a request to check if it is completed or still in progress.

curl https://media.onerouter.pro/v1/audios/tasks/{task_id} \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" 

# For example
curl https://media.onerouter.pro/v1/audios/tasks/f6c274b21ca3472c987e5f6383f3fa7b \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json"

Response with queue task status query

{
  "code": 200,
  "message": "success",
  "data": {
    "task_id": "f6c274b21ca3472c987e5f6383f3fa7b",
    "object": "audio",
    "model": "openai/gpt-4o-mini-tts/text-to-audio",
    "status": "in_progress",
    "fail_reason": "",
    "submit_time": 1780187430,
    "start_time": 1780187430,
    "finish_time": 0,
    "outputs": [],
    "created_at": "2026-05-31 00:30:30"
  }
}

Possible statuses:

  • created
  • in_progress
  • processing
  • completed
  • failed

created -> in_progress → processing → completed | failed

Poll every 1-2 seconds until status is "completed" or "failed"

Get the result (when status is "completed")

Once the request is completed, you can fetch the result. See the Output Schema for the expected result format.

{
  "code": 200,
  "message": "success",
  "data": {
    "task_id": "f6c274b21ca3472c987e5f6383f3fa7b",
    "object": "audio",
    "model": "openai/gpt-4o-mini-tts/text-to-audio",
    "status": "completed",
    "fail_reason": "",
    "submit_time": 1780187430,
    "start_time": 1780187430,
    "finish_time": 1780187443,
    "outputs": [
      "https://storage.googleapis.com/infron_gcs/image%2Faudio_v2%2Foutput%2Ff6c274b21ca3472c987e5f6383f3fa7b_fbd5c084.mp3"
    ],
    "usage": {
      "completion_tokens": 0,
      "output_count": 1,
      "prompt_tokens": 46,
      "total_tokens": 46
    },
    "cost": {
      "total_cost": 2.8e-05,
      "cost_details": {
        "prompt_cost": 0,
        "completion_cost": 0,
        "image_cost": 0,
        "video_cost": 0,
        "audio_cost": 2.8e-05,
        "native_web_search_cost": 0,
        "plugin_web_search_cost": 0,
        "tools_cost": 0,
        "prompt_cache_read_cost": 0,
        "prompt_cache_write_cost": 0,
        "prompt_cache_write_5_min": 0,
        "prompt_cache_write_1_h": 0,
        "reasoning_cost": 0,
        "discount_rate": 1,
        "is_byok": false,
        "byok_cost": 0
      }
    },
    "created_at": "2026-05-31 00:30:30"
  }
}

2. Schema

Input

ParameterTypeRequiredDefaultRangeDescription
modelstringYes--Explore all modelsModel name.
inputstringYes----The text to generate audio for.
voicestringYesalloyalloy, ash, ballad, coral, echo, fable, onyx, nova, sage, shimmer, verseURL of the video to extend.
instructionsstringNo----Control the voice of your generated audio with additional instructions.
response_formatstringNomp3mp3, opus, aac, flac, wav, pcmFormat of the generated audio.
speednumberNo1>=0.25 and <=4.0The speed of the generated audio. Select a value from 0.25 to 4.0.
stream_formatstringNoaudioaudio, sseThe format to stream the audio.

Example request

curl https://media.onerouter.pro/v1/audios/generations \
  -H 'Authorization: Bearer YOUR_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "openai/gpt-4o-mini-tts/text-to-audio",
    "input": "Today is a wonderful day to build with Infron.",
    "voice": "alloy",
    "instructions": "esse incididunt",
    "response_format": "mp3",
    "speed": 1,
    "stream_format": "audio"
  }'

Output

ParameterTypeRangeDescription
codeinteger--Status code of the request.
messagestring--Status message of the request.
dataobject--Data of the task.
data.task_idstring--ID of the task.
data.objectstring--Object type of the task.
data.modelstring--Model name of the task.
data.statusstringcreated, in_progress, processing, completed, failedStatus of the task.
data.fail_reasonstring--Fail reason of the task.
data.submit_timeinteger--Timestamp of the task submission.
data.start_timeinteger--Timestamp of the task start.
data.finish_timeinteger--Timestamp of the task finish.
data.outputsarray--Outputs of the task.
data.created_atstring--Timestamp of the task creation.

Example output

{
  "code": 200,
  "message": "success",
  "data": {
    "task_id": "f6c274b21ca3472c987e5f6383f3fa7b",
    "object": "audio",
    "model": "openai/gpt-4o-mini-tts/text-to-audio",
    "status": "completed",
    "fail_reason": "",
    "submit_time": 1780187430,
    "start_time": 1780187430,
    "finish_time": 1780187443,
    "outputs": [
      "https://storage.googleapis.com/infron_gcs/image%2Faudio_v2%2Foutput%2Ff6c274b21ca3472c987e5f6383f3fa7b_fbd5c084.mp3"
    ],
    "usage": {
      "completion_tokens": 0,
      "output_count": 1,
      "prompt_tokens": 46,
      "total_tokens": 46
    },
    "cost": {
      "total_cost": 2.8e-05,
      "cost_details": {
        "prompt_cost": 0,
        "completion_cost": 0,
        "image_cost": 0,
        "video_cost": 0,
        "audio_cost": 2.8e-05,
        "native_web_search_cost": 0,
        "plugin_web_search_cost": 0,
        "tools_cost": 0,
        "prompt_cache_read_cost": 0,
        "prompt_cache_write_cost": 0,
        "prompt_cache_write_5_min": 0,
        "prompt_cache_write_1_h": 0,
        "reasoning_cost": 0,
        "discount_rate": 1,
        "is_byok": false,
        "byok_cost": 0
      }
    },
    "created_at": "2026-05-31 00:30:30"
  }
}

3. Detailed Pricing

Pricing TypeUnitPrice
Input text1M tokens$0.60
Output audio1M characters$12.00