API Rate Limit Calculator
Plan batch jobs and traffic within an API’s limits. Enter the requests-per-minute limit, the tokens-per-minute limit and tokens per request, the average response time, the number of requests to process and a safety margin. The calculator finds which limit binds, the sustainable requests per minute, the time the batch takes, the number of parallel requests needed and the delay between requests for a single worker.
- Runs in your browser
- No sign-up
- Free to use
Leave empty for APIs without a token limit.
Use exponential backoff with jitter on 429 responses; limits often apply per organisation, model and minute window.
How to use API Rate Limit Calculator
- Enter RPM and TPM limits.
- Enter tokens per request and response time.
- Enter the batch size and safety margin.
- Read throughput, time and concurrency.
API Rate Limit Calculator features
Binding limit
RPM or TPM.
Concurrency
From Little’s law.
Batch time
Minutes or hours.
Safety margin
Stay below limits.
Formula shown
Every result explains how it was calculated.
Private
Runs in your browser; nothing is sent or stored.
When to use API Rate Limit Calculator
- Bulk AI processing jobs.
- Data imports via APIs.
- Load planning for integrations.
- Choosing a rate-limit tier.
API Rate Limit Calculator FAQ
What is Little’s law?
Requests in progress = arrival rate × time each request takes; it gives the concurrency needed.
Why a safety margin?
Limits are often measured over short windows; bursts can trigger 429 errors.
What should I do on a 429 error?
Retry with exponential backoff and jitter.
Do limits apply per key?
Often per organisation and model; check your provider’s documentation.
Working within limits
Rate limits cap how fast you can go, whatever your code does. Knowing which limit binds tells you whether to shorten prompts (tokens) or batch requests (requests).
For large jobs, batch APIs with separate limits and lower prices may be an option.
Limits usually apply per minute, so short bursts above the average rate can still fail even when the hourly total is fine. Spreading requests evenly with a queue, and sizing the number of workers with the concurrency figure shown here, avoids most 429 errors.
Prices, limits and tokenizers differ between providers and change often. Enter the current values from your provider’s pricing and documentation pages, and re-check them before committing to a budget.
Measure real usage once you have it: average tokens per request in production are often different from the estimates used at the planning stage, especially once system prompts, retrieved context and conversation history are included.