Limits into capacity
Plan API throughput and concurrency
Use limits from your dashboard, reserve retry headroom, and validate the estimate against observed p95 latency.
Local capacity planning
Rate Limit & Concurrency Planner
Enter limits from your own provider dashboard. Values are processed only in this page and are not saved. Capacity planning generally benefits from p95 latency rather than the average.
Safe throughput estimate
RPM limit600 req/min
TPM-derived limit300 req/min
- Binding limit
- TPM
- Tokens per request
- 1,000
- Raw effective limit
- 300 req/min
- Safe limit after headroom
- 189 req/min
- Safe requests/second
- 3.15
- Recommended concurrency
- 16
- Burst requests
- 31.5
- Target minus safe limit
- 61 req/min
The target is above the estimated safe limit. This is a planning estimate; implement provider-specific backoff and observe real latency.