Docs

Create an API key in the console, point an OpenAI SDK at https://batchrate.ai/v1, and submit a JSONL batch. Results are ready within the 24 hour window, usually within hours.

Quickstart

Install the official SDK. No batchrate package is required.

pip install openai
from openai import OpenAI

client = OpenAI(
    api_key="br_live_...",
    base_url="https://batchrate.ai/v1",
)

batch_file = client.files.create(file=open("batch.jsonl", "rb"), purpose="batch")
batch = client.batches.create(
    input_file_id=batch_file.id,
    endpoint="/v1/chat/completions",
    completion_window="24h",
)

while batch.status not in ("completed", "failed", "cancelled", "expired"):
    batch = client.batches.retrieve(batch.id)

if batch.output_file_id:
    print(client.files.content(batch.output_file_id).text)

JavaScript is the same idea: new OpenAI({ apiKey, baseURL: "https://batchrate.ai/v1" }).

Local development uses http://127.0.0.1:8788/v1.

Authentication

Send Authorization: Bearer br_live_.... Keys are shown once at creation. batchrate stores only a SHA-256 hash. Revoked keys stop working immediately. The console can also call these routes with the session cookie; cross-site cookie posts are rejected.

Files

POST /v1/files

Multipart form with purpose=batch and file. The body must be valid batch JSONL within the configured size and line limits. Response:

{
  "id": "file-...",
  "object": "file",
  "bytes": 120,
  "created_at": 1710000000,
  "filename": "batch.jsonl",
  "purpose": "batch",
  "status": "processed",
  "status_details": null
}

GET /v1/files/{id}

Returns the same file object if you own it.

GET /v1/files/{id}/content

Returns the raw bytes. Use this for input, output, and error files.

DELETE /v1/files/{id}

Deletes a file you own and removes the object from storage immediately.

{
  "id": "file-...",
  "object": "file",
  "deleted": true
}

If the input file is still used by a batch in validating, in_progress, finalizing, or cancelling, the API returns file_in_use and names that batch. Cancel the batch first. A batch that is only cancelling because work is still leased stays blocked until it reaches a terminal status. Output and error files can be deleted at any time. The batch row stays so the ledger can still point at the job. A deleted file returns file_not_found on later reads.

File retention

Input files are deleted 7 days after every batch that used them reaches a terminal state (completed, cancelled, failed, or expired). An input file that is never used is deleted 7 days after upload. Output and error files are deleted 30 days after they are created. The periods are INPUT_RETENTION_DAYS and OUTPUT_RETENTION_DAYS in app_config. You can delete sooner with DELETE /v1/files/{id} or the console. Automatic deletion removes the stored object and marks the database row deleted. Ledger rows stay for accounting.

Batches

POST /v1/batches

{
  "input_file_id": "file-...",
  "endpoint": "/v1/chat/completions",
  "completion_window": "24h",
  "metadata": { "dataset": "support-queue" }
}

endpoint must be /v1/chat/completions. completion_window must be 24h. The batch moves to in_progress after validation. If the available credit balance is below the estimate, the API returns insufficient_quota.

GET /v1/files/{id}/estimate

Recomputes a live cost estimate for an uploaded file. output_is_estimate is always true. Output tokens come from each line's max_tokens (or max_completion_tokens), otherwise default_output_tokens, capped by max_output_tokens_estimate. warning is set when free tokens plus the available balance will not cover price_microusd. The same object is included as estimate on the file resource. A created batch stores the quote used for its hold, with output tokens still marked as an estimate.

GET /v1/batches and GET /v1/batches/{id}

List responses use object: "list", data, first_id, last_id, and has_more. Pass limit and after to page. A batch includes request_counts, output_file_id, error_file_id, usage, and estimate.

POST /v1/batches/{id}/cancel

Queued lines are cancelled. Lines already leased on a GPU can still finish. When nothing is left in flight, the batch becomes cancelled and any completed lines are still billed and downloadable.

Statuses: validating, in_progress, finalizing, completed, failed, cancelling, cancelled, expired.

Models

GET /v1/models

Lists enabled models. The first one is qwen3.6-35b-a3b (Qwen3.6-35B-A3B FP8). Prices are not on this route. Read /api/pricing or the pricing page.

JSONL format

{"custom_id":"req-1","method":"POST","url":"/v1/chat/completions","body":{"model":"qwen3.6-35b-a3b","messages":[{"role":"user","content":"Summarize this ticket."}]}}

Output lines follow the OpenAI batch result shape:

{"id":"batch_req_...","custom_id":"req-1","response":{"status_code":200,"request_id":"req_...","body":{"id":"chatcmpl-...","object":"chat.completion","choices":[{"message":{"role":"assistant","content":"..."}}],"usage":{"prompt_tokens":10,"completion_tokens":20,"total_tokens":30}}},"error":null}

Failed lines go to the error file with response: null and an error object of code and message.

Errors

{"error":{"message":"...","type":"invalid_request_error","param":null,"code":"invalid_file"}}

Common codes: invalid_api_key, invalid_file, insufficient_quota, rate_limit_exceeded, file_not_found, file_in_use, batch_not_found.

Billing

Credits are prepaid. Checkout is Stripe-hosted. The success page does not add credits; the signed webhook does, and only when payment_status is paid. Duplicate events are ignored. Token prices are micro-USD per million tokens in model_prices. The launch rate for qwen3.6-35b-a3b is stored there ($0.05 / M input, $0.40 / M output) with is_placeholder = 0. min_topup_cents is 1000 ($10) and default_topup_cents is 2500 ($25). Verified accounts also receive FREE_TOKENS_ON_SIGNUP tokens once. Those tokens are spent before the dollar balance. A public calculator at /pricing#estimator posts to POST /api/estimate and shows factual price rows from reference_prices, each marked as a reference price with its as-of date (2026-10-10 at launch).