All guides

Set Request and Token Rate Limits

Size limits from real traffic, give every caller its own allowance, and give callers a clear signal to retry.

9 min read Updated October 2026LLM FinOps
On this page
  1. 01What you'll set up
  2. 02Two limits, two protections
  3. 03Size the limits from real traffic
  4. 04Set the limits
  5. 05Every caller gets its own allowance
  6. 06What callers see
  7. 07Retry the right way
  8. 08Put limits next to the other caps
  9. 09Test the limit
  10. 10Keep it as code
  11. 11Roll out in steps
  12. 12If something goes wrong

01

What you'll set up

Rate limits keep traffic smooth. In about nine minutes you will size a request limit and a token limit per minute from real traffic, set them, test them, and make sure your callers handle the answer well.

  • A request limit and a token limit per minute
  • Numbers based on your own traffic, not guesses
  • A separate allowance for every caller
  • A readable response and a retry pattern for developers

You need an owner or admin role and a policy bound to your traffic.

02

Two limits, two protections

Use them together. They guard against different problems.

LimitWhat it countsProtects against
Request rate limit / minRequests per minuteA retry storm or a runaway script
Token rate limit / minThe tokens requests can use each minute, the prompt plus the output limitA few very large requests that use up a provider quota

03

Size the limits from real traffic

Start from what the app does at its busiest, then add headroom.

  1. 1

    Find a busy hour

    In AI → Explorer, group by App and choose 24h. Use Cost Trend to spot the busiest hour, and Spend Breakdown for the day's requests and tokens.

  2. 2

    Turn it into a per-minute number

    Divide the requests in that hour by 60, then double the result for bursts.

  3. 3

    Size the token limit

    Multiply that per-minute request peak by the average tokens a request uses.

  4. 4

    Add headroom

    Round up. A limit that sits at your peak blocks real work.

MeasureValue
Requests in the busiest hour9,000
Average per minute150
Request rate limit / min (double, rounded up)400
Average tokens per request1,200
Token rate limit / min (400 × 1,200, rounded up)500,000
Example, for illustration

04

Set the limits

Rate limits are on the Advanced step of the policy form.

  1. 1

    Open the policy

    Go to AI → Policies and open the policy that governs your traffic.

  2. 2

    Open Throughput limits

    On Advanced, find the Throughput limits section.

  3. 3

    Enter the limits

    Fill in Request rate limit / min and Token rate limit / min. Leave a field empty for no limit.

  4. 4

    Set Enforcement mode to Enforce

    Rate limits take effect on an enforcing policy. Switch one app's policy first, then widen.

  5. 5

    Save

    The limits apply to new requests at once.

AI → Policies → Create policy → Advanced

Throughput limits

3Request rate limit / min
600
3Token rate limit / min
500000
Set both limits, then switch the policy to Enforce.

05

Every caller gets its own allowance

Limits are counted per minute and kept separate for each caller. A noisy app uses up its own allowance and leaves every other app on the policy untouched.

CallerRequests this minuteResult
support-assistant90Allowed, 30 left
nightly-batch14020 requests over the limit return 429
search-service40Unaffected by the batch job
A policy with a limit of 120 requests per minute

06

What callers see

A request over a limit returns HTTP 429. It names which limit was reached and asks the caller to retry shortly. Limits are counted in one-minute windows, so a short wait is usually enough.

HTTP 429
{
  "error": "Your AI request was rate limited by the active Cloptima policy. Please retry shortly.",
  "reason": "request_rate_limit_exceeded",
  "violations": ["request_rate_limit_exceeded"]
}
ReasonLimit reached
request_rate_limit_exceededRequest rate limit / min
token_rate_limit_exceededToken rate limit / min

07

Retry the right way

A good retry turns a limit into a short pause instead of an error.

  • Retry only on HTTP 429, never on a 402 or 403
  • Wait, then retry, and back off further each time
  • Add a little random jitter so many workers do not retry together
  • Queue batch work instead of sending it all at once
python
import random
import time
from openai import OpenAI, APIStatusError

client = OpenAI(base_url="https://api.cloptima.ai/v1/ai", api_key="clop_vk_...")

def ask(prompt, attempts=4):
    for attempt in range(attempts):
        try:
            return client.chat.completions.create(
                model="gpt-4o-mini",
                messages=[{"role": "user", "content": prompt}],
            )
        except APIStatusError as error:
            if error.status_code != 429 or attempt == attempts - 1:
                raise
            time.sleep(2 ** attempt + random.random())  # back off, with jitter

08

Put limits next to the other caps

Rate limits are one of several caps on the same policy. Each stops a different failure.

CapWhereStops
Request and token rate limitsAdvanced → Throughput limitsBursts in a minute
Daily and monthly budgetBasics and AdvancedSpend over a day or a month
Maximum input and output sizeAdvancedOne oversized request
Tool-call and loop capsAdvanced → Agent-loop & execution capsAn agent that loops on itself

09

Test the limit

Prove it with a throwaway policy and a low limit before you set real numbers.

  1. 1

    Set a low limit

    On a test policy, set Request rate limit / min to 10 and Mode to Enforce. Bind it to a test key.

  2. 2

    Send a burst

    Send 30 requests in under a minute.

  3. 3

    Read the result

    The first requests succeed. Later ones return 429 with request_rate_limit_exceeded.

bash
for i in $(seq 1 30); do
  curl -s -o /dev/null -w "%{http_code}\n" https://api.cloptima.ai/v1/ai/chat/completions \
    -H "Authorization: Bearer $CLOPTIMA_VIRTUAL_KEY" \
    -H "Content-Type: application/json" \
    -d '{"model": "gpt-4o-mini", "max_tokens": 16, "messages": [{"role": "user", "content": "Say hello"}]}'
done

10

Keep it as code

Limits can live in Terraform with the rest of the policy.

main.tf
resource "cloptima_llm_gateway_policy" "support_production" {
  name                          = "support-production"
  mode                          = "enforce"
  request_rate_limit_per_minute = 400
  token_rate_limit_per_minute   = 500000
}

11

Roll out in steps

Add limits the same way you add any control: one app first.

  1. 1

    Start with the Request Simulator

    Check that the right policy applies to the app's key.

  2. 2

    Enforce on one app

    Set the limits on a policy bound to a single app and switch it to Enforce.

  3. 3

    Watch the Blocked column

    In the Explorer, a few 429s at your peak mean the limit is close. Raise it if real work is waiting.

  4. 4

    Widen the binding

    Bind the policy to more apps.

12

If something goes wrong

Most problems come down to the value or the mode.

What you seeLikely causeFix
429 while traffic looks lowThe limit is below your real peakRaise the limit
429 on token limits onlyRequests allow large outputsSet a realistic maximum output size, or raise the token limit
Limits never triggerThe policy is in Monitor only modeSwitch Mode to Enforce
Only some apps are limitedEach caller is counted separatelyCheck which apps the policy is bound to

Put This Guide Into Practice

Cloptima automates the strategies described in this guide.

No credit card required
5-minute setup
Free trial