All guides

Fix Blocked AI Gateway Requests

Read the gateway's block response, find the rule that applied, and fix the request or the policy.

7 min read Updated October 2026LLM FinOps
On this page
  1. 01Anatomy of a block
  2. 02Start from the HTTP status
  3. 03Access and model rules
  4. 04Size and agent limits
  5. 05Rate and budget limits
  6. 06Guardrail blocks
  7. 07Missing attribution
  8. 08Handle blocks in your app
  9. 09Find the rule in the console

01

Anatomy of a block

A blocked request returns a JSON body with a readable message, a reason code, and the list of rules it broke. Some responses add details, such as the size you sent and the limit that applies.

HTTP 403
{
  "error": "Your AI request was blocked because the estimated input (18200 tokens) exceeds the active Cloptima policy limit (8000 tokens).",
  "reason": "max_input_tokens_exceeded",
  "violations": ["max_input_tokens_exceeded"],
  "details": { "estimated_input_tokens": 18200, "max_input_tokens": 8000 }
}
FieldWhat it holds
errorA sentence you can show to a developer
reasonA stable code to branch on in your code
violationsEvery rule the request broke
detailsNumbers that explain the block, when there are any

02

Start from the HTTP status

The status code tells you which kind of rule applied. Start there, then read the reason.

StatusMeaningGo to
400The request is missing required labelsMissing attribution
402A budget or credit limit was reachedRate and budget limits
403A policy, tool, size, or guardrail rule blocked itAccess, size, and guardrail sections
429A rate limit was reachedRate and budget limits
502 or 503A gateway check could not completeRetry shortly. For guardrail and rate checks, your policy's fail mode decides whether the request stops or continues.

03

Access and model rules

These mean the policy does not allow what the request asked for.

ReasonWhat it meansFix
policy_not_configuredNo policy is bound to this key, app, or teamBind a policy
provider_not_allowedThe provider is not on the allow listAllow it, or choose another provider
model_not_allowedThe model is not on the allow listAllow it, or use an allowed model
model_deniedThe model is on the deny listUse another model, or remove it from the deny list
tool_not_allowed, tool_deniedThe tool is not allowedAllow the tool on the policy
tool_server_not_allowed and relatedThe tool server is not allowedAllow the tool server on the policy
required_metadata_missingA label the policy requires was not sentSend the required team and app labels

04

Size and agent limits

These mean the request is larger or longer-running than the policy allows.

ReasonWhat it meansFix
max_input_tokens_exceededThe prompt is larger than allowedSend less, or raise the limit
max_output_tokens_exceededThe requested answer length is above the limitRequest a shorter answer, or raise the limit
max_tool_calls_exceededThe agent has made more tool calls than allowedRaise the limit if the behavior is expected
max_retry_count_exceeded, max_loop_iterations_exceededThe agent has retried or looped too oftenFix the loop, or raise the limit

05

Rate and budget limits

Rate limits return HTTP 429. Budget and credit blocks return HTTP 402.

ReasonWhat it meansFix
request_rate_limit_exceededToo many requests this minuteRetry shortly, or raise the limit
token_rate_limit_exceededToo many tokens this minuteRetry shortly, or raise the limit
A budget or credit block (402)The budget or credits are used upRaise the budget, wait for the next period, or add credits

06

Guardrail blocks

A guardrail block returns HTTP 403 with the reason gateway_guardrail_blocked.

The violations list names the rule that matched, for example a built-in secret rule such as github_token or one of your own custom rules, which appear as custom_ followed by the rule id. The response never repeats the matched text.

07

Missing attribution

If your setup requires team and app labels and a request has neither, the gateway returns HTTP 400.

Set the team and app on the virtual key, or send them as the x-cloptima-team and x-cloptima-app headers.

08

Handle blocks in your app

Branch on the reason code, not on the message text. The message is for people. The code is stable.

python
from openai import OpenAI, APIStatusError

client = OpenAI(base_url="https://api.cloptima.ai/v1/ai", api_key="clop_vk_...")

try:
    reply = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": "Summarize this ticket"}],
    )
except APIStatusError as error:
    body = error.response.json()
    if body.get("reason") == "request_rate_limit_exceeded":
        ...  # back off and retry
    else:
        print(body["error"])  # show the readable message to the developer

09

Find the rule in the console

Blocked requests appear in the Policy Violations card in the Audit tab.

  1. 1

    Open Audit

    Find the blocked request by time and app.

  2. 2

    Match the reason to a setting

    The tables above map each reason to the field you would change.

  3. 3

    Confirm the fix in the Request Simulator

    Check the change before you save it for everyone.

Put This Guide Into Practice

Cloptima automates the strategies described in this guide.

No credit card required
5-minute setup
Free trial