/run traffic, so submitting a batch never delays your interactive requests.
When to use batch vs /run
Choose batch when your workload can tolerate multi-hour latency — for example, nightly dataset processing, pre-computing embeddings, or running evaluations.
Batch lifecycle
A batch moves through the following states:- OPEN — The batch is a draft. You can add, update, or remove individual requests. Batch workers have not started any work.
- FINALIZED — The batch is locked. No further requests can be added or removed. The batch is now eligible for execution and will be picked up by batch workers.
- RUNNING — At least one request in the batch is being processed by a worker.
- COMPLETED — All requests have reached a terminal state (completed or failed).
- FAILED — The batch itself failed before or during execution (distinct from individual request failures within an otherwise-completed batch).
- CANCELLED — You cancelled the batch. See Cancellation for details.
/finalize before the batch begins processing. An OPEN batch will not be executed.
API walkthrough
1. Create a batch
requests array uses the same shape as a standard /run call — a JSON object with an input field.
2. Add more requests
While the batch is OPEN, append additional requests:3. Finalize the batch
Once you’ve added all requests, finalize the batch to make it eligible for execution:FINALIZED and requests are locked. You can no longer add or remove individual requests.
4. Poll batch status
Check overall progress by fetching the batch summary:status is COMPLETED, FAILED, or CANCELLED, the batch has reached a terminal state.
5. Retrieve results
Fetch paginated results for all child requests in the batch:nextCursor as a query parameter to retrieve the next page.
Full API reference
For full request and response schemas, see the API reference.
Monitoring batches in the console
Open your endpoint in the Runpod console and select the Batch tab to see all batches. Each row shows the batch name, status, and progress counts. Click a batch to open the detail view, which shows:- Top-level status and progress
- Per-request rows with status, timestamps, and error messages for failed requests
- Links to the full request detail view for each child request
Notifications
When a batch reaches a terminal state (COMPLETED, FAILED, or CANCELLED), Runpod sends:- Console Inbox notification — includes batch ID, endpoint name, terminal status, and item counts (completed / failed / total)
- Webhook event — if your account has a webhook subscription configured for batch events
Cancellation
To cancel a batch:- Queued requests are cancelled immediately and are not billed.
- In-progress requests are allowed to finish and are billed normally.
CANCELLED once all in-progress work has drained.
Limits
If you need to exceed these limits, contact support.
Billing
Batch jobs are billed at the same rate as standard serverless requests on your endpoint. For enterprise customers, flex worker discounts apply to batch jobs. Billing is based on the compute time used by each child request, regardless of whether the batch was later cancelled (in-progress requests that completed before cancellation are billed normally).Error handling
Individual request failures — A failed child request does not fail the entire batch. The batch continues processing remaining requests and reaches COMPLETED status. Inspect failed requests via the console or theGET .../requests endpoint; each failed request includes an error message from the handler.
Batch-level failure — If the batch itself fails (status FAILED), it indicates a systemic problem rather than individual handler errors. Contact support if you see this state and cannot explain it from request-level errors.
Redis durability — Batch jobs use the same Redis-backed queue as standard serverless requests. In the event of a Redis failure, queued batch requests may be lost. This is an MVP limitation that applies equally to /run traffic.
Known limitations
- Batch jobs inherit the GPU type configured on your endpoint. You cannot specify a different GPU per batch or per request.
- There is no per-request scheduling or ordering. Requests within a batch are processed in an unspecified order.
- Cost estimation before finalization is not available at launch.
- Batch workers are scheduled based on global queue urgency and off-peak capacity. Start time within the 24h SLA is not guaranteed.