On this page
01
What you'll set up
Rate limits keep traffic smooth. In about nine minutes you will size a request limit and a token limit per minute from real traffic, set them, test them, and make sure your callers handle the answer well.
- A request limit and a token limit per minute
- Numbers based on your own traffic, not guesses
- A separate allowance for every caller
- A readable response and a retry pattern for developers
You need an owner or admin role and a policy bound to your traffic.
02
Two limits, two protections
Use them together. They guard against different problems.
| Limit | What it counts | Protects against |
|---|---|---|
| Request rate limit / min | Requests per minute | A retry storm or a runaway script |
| Token rate limit / min | The tokens requests can use each minute, the prompt plus the output limit | A few very large requests that use up a provider quota |
03
Size the limits from real traffic
Start from what the app does at its busiest, then add headroom.
- 1
Find a busy hour
In AI → Explorer, group by App and choose 24h. Use Cost Trend to spot the busiest hour, and Spend Breakdown for the day's requests and tokens.
- 2
Turn it into a per-minute number
Divide the requests in that hour by 60, then double the result for bursts.
- 3
Size the token limit
Multiply that per-minute request peak by the average tokens a request uses.
- 4
Add headroom
Round up. A limit that sits at your peak blocks real work.
| Measure | Value |
|---|---|
| Requests in the busiest hour | 9,000 |
| Average per minute | 150 |
| Request rate limit / min (double, rounded up) | 400 |
| Average tokens per request | 1,200 |
| Token rate limit / min (400 × 1,200, rounded up) | 500,000 |
04
Set the limits
Rate limits are on the Advanced step of the policy form.
- 1
Open the policy
Go to AI → Policies and open the policy that governs your traffic.
- 2
Open Throughput limits
On Advanced, find the Throughput limits section.
- 3
Enter the limits
Fill in Request rate limit / min and Token rate limit / min. Leave a field empty for no limit.
- 4
Set Enforcement mode to Enforce
Rate limits take effect on an enforcing policy. Switch one app's policy first, then widen.
- 5
Save
The limits apply to new requests at once.
Throughput limits
- 3Request rate limit / min
- 600
- 3Token rate limit / min
- 500000
05
Every caller gets its own allowance
Limits are counted per minute and kept separate for each caller. A noisy app uses up its own allowance and leaves every other app on the policy untouched.
| Caller | Requests this minute | Result |
|---|---|---|
| support-assistant | 90 | Allowed, 30 left |
| nightly-batch | 140 | 20 requests over the limit return 429 |
| search-service | 40 | Unaffected by the batch job |
06
What callers see
A request over a limit returns HTTP 429. It names which limit was reached and asks the caller to retry shortly. Limits are counted in one-minute windows, so a short wait is usually enough.
{
"error": "Your AI request was rate limited by the active Cloptima policy. Please retry shortly.",
"reason": "request_rate_limit_exceeded",
"violations": ["request_rate_limit_exceeded"]
}| Reason | Limit reached |
|---|---|
| request_rate_limit_exceeded | Request rate limit / min |
| token_rate_limit_exceeded | Token rate limit / min |
07
Retry the right way
A good retry turns a limit into a short pause instead of an error.
- Retry only on HTTP 429, never on a 402 or 403
- Wait, then retry, and back off further each time
- Add a little random jitter so many workers do not retry together
- Queue batch work instead of sending it all at once
import random
import time
from openai import OpenAI, APIStatusError
client = OpenAI(base_url="https://api.cloptima.ai/v1/ai", api_key="clop_vk_...")
def ask(prompt, attempts=4):
for attempt in range(attempts):
try:
return client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
)
except APIStatusError as error:
if error.status_code != 429 or attempt == attempts - 1:
raise
time.sleep(2 ** attempt + random.random()) # back off, with jitter08
Put limits next to the other caps
Rate limits are one of several caps on the same policy. Each stops a different failure.
| Cap | Where | Stops |
|---|---|---|
| Request and token rate limits | Advanced → Throughput limits | Bursts in a minute |
| Daily and monthly budget | Basics and Advanced | Spend over a day or a month |
| Maximum input and output size | Advanced | One oversized request |
| Tool-call and loop caps | Advanced → Agent-loop & execution caps | An agent that loops on itself |
09
Test the limit
Prove it with a throwaway policy and a low limit before you set real numbers.
- 1
Set a low limit
On a test policy, set Request rate limit / min to 10 and Mode to Enforce. Bind it to a test key.
- 2
Send a burst
Send 30 requests in under a minute.
- 3
Read the result
The first requests succeed. Later ones return 429 with request_rate_limit_exceeded.
for i in $(seq 1 30); do
curl -s -o /dev/null -w "%{http_code}\n" https://api.cloptima.ai/v1/ai/chat/completions \
-H "Authorization: Bearer $CLOPTIMA_VIRTUAL_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-4o-mini", "max_tokens": 16, "messages": [{"role": "user", "content": "Say hello"}]}'
done10
Keep it as code
Limits can live in Terraform with the rest of the policy.
resource "cloptima_llm_gateway_policy" "support_production" {
name = "support-production"
mode = "enforce"
request_rate_limit_per_minute = 400
token_rate_limit_per_minute = 500000
}11
Roll out in steps
Add limits the same way you add any control: one app first.
- 1
Start with the Request Simulator
Check that the right policy applies to the app's key.
- 2
Enforce on one app
Set the limits on a policy bound to a single app and switch it to Enforce.
- 3
Watch the Blocked column
In the Explorer, a few 429s at your peak mean the limit is close. Raise it if real work is waiting.
- 4
Widen the binding
Bind the policy to more apps.
12
If something goes wrong
Most problems come down to the value or the mode.
| What you see | Likely cause | Fix |
|---|---|---|
| 429 while traffic looks low | The limit is below your real peak | Raise the limit |
| 429 on token limits only | Requests allow large outputs | Set a realistic maximum output size, or raise the token limit |
| Limits never trigger | The policy is in Monitor only mode | Switch Mode to Enforce |
| Only some apps are limited | Each caller is counted separately | Check which apps the policy is bound to |