> For the complete documentation index, see [llms.txt](https://infronai.gitbook.io/docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://infronai.gitbook.io/docs/overview/quickstart/batch.md).

# Batch

Infron Batch runs inference offline: submit multiple requests, receive a job ID, and download the results when processing finishes. Each job uses one model.

**Endpoint:** `https://llm.onerouter.pro/v1/batch`&#x20;

### 1. Set up your client

Install the Python dependency and set your API key and a model enabled for Infron Batch:

```bash
pip install requests
export INFRON_API_KEY='YOUR_API_KEY'
export INFRON_BATCH_MODEL='infron-batch/deepseek-v4.1-flash'
```

The example uses `infron-batch/deepseek-v4.1-flash` from the reference guide. Confirm that your key can access it, or replace it with an [available Batch model ID](https://infron.ai/models?tab=batch).

Run the following Python snippets in order, sharing this configuration:

```python
import os
import requests

BASE_URL = "https://llm.onerouter.pro"
MODEL = os.environ["INFRON_BATCH_MODEL"]
HEADERS = {
    "Authorization": f"Bearer {os.environ['INFRON_API_KEY']}",
    "Content-Type": "application/json",
    "User-Agent": "infron-batch-guide/1.0",
    "Accept": "application/json",
    "Connection": "keep-alive",
}
```

### 2. Create a Batch job

**`POST /v1/batch`**

Set the model at the top level and put individual requests in `requests`. Use a unique `custom_id` to identify each business record. Each `body` uses Chat Completions format. This example submits 11 requests, with IDs `q1` through `q11`, using English prompts.

```python
payload = {
    "model": MODEL,
    "requests": [
        {
            "custom_id": "q1",
            "body": {
                "messages": [{"role": "user", "content": "Describe Beijing in one sentence."}],
            },
        },
        {
            "custom_id": "q2",
            "body": {
                "messages": [{"role": "user", "content": "Describe Shanghai in one sentence."}],
            },
        },
        {
            "custom_id": "q3",
            "body": {
                "messages": [{"role": "user", "content": "Describe Shanghai in one sentence."}],
            },
        },
        {
            "custom_id": "q4",
            "body": {
                "messages": [{"role": "user", "content": "Describe Shanghai in one sentence."}],
            },
        },
        {
            "custom_id": "q5",
            "body": {
                "messages": [{"role": "user", "content": "Describe Shanghai in one sentence."}],
            },
        },
        {
            "custom_id": "q6",
            "body": {
                "messages": [{"role": "user", "content": "Describe Shanghai in one sentence."}],
            },
        },
        {
            "custom_id": "q7",
            "body": {
                "messages": [{"role": "user", "content": "Describe Shanghai in one sentence."}],
            },
        },
        {
            "custom_id": "q8",
            "body": {
                "messages": [{"role": "user", "content": "Describe Shanghai in one sentence."}],
            },
        },
        {
            "custom_id": "q9",
            "body": {
                "messages": [{"role": "user", "content": "Describe Shanghai in one sentence."}],
            },
        },
        {
            "custom_id": "q10",
            "body": {
                "messages": [{"role": "user", "content": "Describe Shanghai in one sentence."}],
            },
        },
        {
            "custom_id": "q11",
            "body": {
                "messages": [{"role": "user", "content": "Describe Shanghai in one sentence."}],
            },
        },
    ],
}

response = requests.post(
    f"{BASE_URL}/v1/batch",
    headers=HEADERS,
    json=payload,
    timeout=180,
)
response.raise_for_status()
batch_id = response.json()["id"]
print(batch_id)
```

Success returns **`202 Accepted`** and a complete job object. The queued record from the reference guide below shows every field; actual IDs, quota amounts, and timestamps depend on your request:

```json
{
  "id": "batch_f32a93f000de4d959c907f167a94137b",
  "model": "infron-batch/deepseek-v4.1-flash",
  "status": "queued",
  "input_bytes": 1036,
  "request_count": 11,
  "completed_count": 0,
  "failed_count": 0,
  "estimated_quota": 4994,
  "charged_quota": 0,
  "expires_at": 1791535183,
  "result_bytes": 0,
  "s3_usage": {
    "put": 1,
    "put_bytes": 1036
  },
  "created_at": 1791448782,
  "updated_at": 1791448800
}
```

> Put generation parameters such as `temperature` and `max_tokens` inside each `body`. A job uses one model and produces non-streaming results.

### 3. Check job progress

**`GET /v1/batch/{id}`**

```python
response = requests.get(
    f"{BASE_URL}/v1/batch/{batch_id}",
    headers=HEADERS,
    timeout=30,
)
response.raise_for_status()
job = response.json()
print(job)
```

Complete response after processing finishes, preserving every field and numeric example value from the reference guide. The signed URL contains illustrative signature placeholders and cannot be used for a download:

```json
{
  "id": "batch_f32a93f000de4d959c907f167a94137b",
  "model": "infron-batch/deepseek-v4.1-flash",
  "status": "completed",
  "input_bytes": 1036,
  "request_count": 11,
  "completed_count": 11,
  "failed_count": 0,
  "estimated_quota": 4994,
  "charged_quota": 721,
  "expires_at": 1791535183,
  "result_bytes": 3005,
  "s3_usage": {
    "get": 14,
    "put": 13,
    "get_bytes": 19423,
    "put_bytes": 20356
  },
  "created_at": 1791448782,
  "updated_at": 1791448849,
  "completed_at": 1791448843,
  "result_url": "https://infron-batch-prod.s3.us-west-2.amazonaws.com/infron-batch/batch_f32a93f000de4d959c907f167a94137b/output-1.jsonl.gz?X-Amz-Algorithm=...&X-Amz-Expires=900&X-Amz-Signature=<SIGNATURE_OMITTED>",
  "result_expires_in": 900
}
```

| Field                                      | Meaning                                                                                                                   |
| ------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------- |
| `id`, `model`                              | Job ID and selected model                                                                                                 |
| `status`                                   | `queued`: waiting or executing; `finalizing`: assembling results; `completed`, `cancelled`, and `failed`: terminal states |
| `input_bytes`                              | Input size in bytes                                                                                                       |
| `request_count`                            | Total requests                                                                                                            |
| `completed_count`, `failed_count`          | Successful and unsuccessful counts; unsuccessful includes failures, cancellations, and expirations                        |
| `estimated_quota`, `charged_quota`         | Estimated and settled quota; the estimate is not a spending cap                                                           |
| `expires_at`                               | Queue deadline in Unix seconds, not result-file expiration                                                                |
| `result_bytes`                             | Compressed result-file size in bytes                                                                                      |
| `s3_usage`                                 | Storage statistics: `get` / `put` count operations; `get_bytes` / `put_bytes` count bytes                                 |
| `created_at`, `updated_at`, `completed_at` | Creation, update, and termination times in Unix seconds; `completed_at` is absent before termination                      |
| `result_url`, `result_expires_in`          | Individual-job query download URL and lifetime in seconds, when results are ready and within retention                    |
| `cancelled_at`                             | Cancellation-request time in Unix seconds, present after cancellation is requested                                        |
| `error_code`, `error_message`              | Job-level error code and message, when present                                                                            |

Poll about every 10 seconds. On `429`, wait according to `Retry-After`. **`completed` means the job ended, not that every request succeeded.** For `failed`, inspect `error_code` and `error_message`.

### 4. Download and read results

The output is gzip-compressed JSONL with one record per request. Match records to your input using `custom_id`.

```python
import gzip
import json

# Use the query response's URL. Do not attach the Infron API key to storage calls.
with requests.get(
    job["result_url"],
    headers={
        "User-Agent": "infron-batch-guide/1.0",
        "Accept": "application/octet-stream",
        "Accept-Encoding": "identity",
        "Connection": "keep-alive",
    },
    stream=True,
    timeout=180,
) as response:
    response.raise_for_status()
    with open("output.jsonl.gz", "wb") as output:
        for chunk in response.raw.stream(1024 * 1024, decode_content=False):
            output.write(chunk)

with gzip.open("output.jsonl.gz", "rt", encoding="utf-8") as result_file:
    for line in result_file:
        item = json.loads(line)
        if item.get("error"):
            print(item["custom_id"], "Error:", item["error"])
        else:
            print(item["custom_id"], item["response"]["body"])
```

Example successful record preserving every field shown in the reference guide. Values inside `choices` and `usage`, and the `request_id`, are illustrative:

```json
{
  "completed_at": 1759900000,
  "index": 0,
  "custom_id": "q1",
  "response": {
    "status_code": 200,
    "body": {
      "choices": [
        {
          "index": 0,
          "message": {
            "role": "assistant",
            "content": "Beijing is the capital of China."
          },
          "finish_reason": "stop"
        }
      ],
      "usage": {
        "prompt_tokens": 16,
        "completion_tokens": 8,
        "total_tokens": 24
      },
      "request_id": "request_example",
      "service_tier": "standard"
    }
  },
  "error": null
}
```

Example failed record:

```json
{
  "index": 1,
  "custom_id": "q2",
  "response": null,
  "error": {
    "code": "insufficient_quota",
    "message": "Batch request failed: insufficient_quota"
  }
}
```

Download URLs default to 15 minutes of validity. Query the job again if the URL expires. Result files default to 7 days of retention; production configuration may vary, so download promptly.

### 5. Other operations

#### List jobs

**`GET /v1/batch`**

```python
response = requests.get(f"{BASE_URL}/v1/batch", headers=HEADERS, timeout=30)
response.raise_for_status()
page = response.json()

# Pass the returned next_cursor unchanged to after.
if page.get("next_cursor"):
    response = requests.get(
        f"{BASE_URL}/v1/batch",
        headers=HEADERS,
        params={"after": page["next_cursor"]},
        timeout=30,
    )
    response.raise_for_status()
```

Complete list response example:

```json
{
  "object": "list",
  "data": [
    {
      "id": "batch_f32a93f000de4d959c907f167a94137b",
      "model": "infron-batch/deepseek-v4.1-flash",
      "status": "completed",
      "input_bytes": 1036,
      "request_count": 11,
      "completed_count": 11,
      "failed_count": 0,
      "estimated_quota": 4994,
      "charged_quota": 721,
      "expires_at": 1791535183,
      "result_bytes": 3005,
      "s3_usage": {
        "get": 14,
        "put": 13,
        "get_bytes": 19423,
        "put_bytes": 20356
      },
      "created_at": 1791448782,
      "updated_at": 1791448849,
      "completed_at": 1791448843
    }
  ],
  "next_cursor": ""
}
```

Each page contains up to 50 jobs. An empty `next_cursor` means there is no next page. Do not substitute a `batch_...` ID for the cursor. List responses do not include download URLs; query an individual job to obtain `result_url`.

#### Cancel a job

**`POST /v1/batch/{id}/cancel`**

```python
response = requests.post(
    f"{BASE_URL}/v1/batch/{batch_id}/cancel",
    headers=HEADERS,
    timeout=30,
)
response.raise_for_status()
print(response.json())
```

Success returns `200` and a complete job object. This cancellation response illustrates the structure; the cancellation timestamp is an example value:

```json
{
  "id": "batch_f32a93f000de4d959c907f167a94137b",
  "model": "infron-batch/deepseek-v4.1-flash",
  "status": "cancelling",
  "input_bytes": 1036,
  "request_count": 11,
  "completed_count": 0,
  "failed_count": 0,
  "estimated_quota": 4994,
  "charged_quota": 0,
  "expires_at": 1791535183,
  "result_bytes": 0,
  "s3_usage": {
    "put": 1,
    "put_bytes": 1036
  },
  "created_at": 1791448782,
  "updated_at": 1791448801,
  "cancelled_at": 1791448801
}
```

Continue polling until a terminal state. Repeated cancellation is safe; ended jobs return their current state. Completed requests may retain results and incur charges.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://infronai.gitbook.io/docs/overview/quickstart/batch.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `automate deployments from our CI pipeline` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
