All guides

Create Your First AI Gateway Policy

Set model access, limits, and spend controls for AI requests, test them without live traffic, then switch to Enforce.

10 min read Updated October 2026LLM FinOps
On this page
  1. 01What you'll set up
  2. 02How a policy fits together
  3. 03Create the policy
  4. 04Choose what is allowed
  5. 05Add limits that match the app
  6. 06Set a budget
  7. 07Decide how failures behave
  8. 08Bind it and send a test request
  9. 09Switch to Enforce and break it on purpose
  10. 10If something goes wrong

01

What you'll set up

In about ten minutes you will have a policy that limits which models an app can use, how large its requests can be, and how much it can spend per day. You will test it before any live traffic is affected.

  • A policy in Monitor only mode, so nothing is blocked while you tune it
  • Model access, size limits, and a rate limit that fit your app
  • A daily budget with the enforcement style you choose
  • A tested path from a simulated request to a real one

You will need an owner or admin role in your organization, an AI plan (policy enforcement is part of the AI plans), and a virtual key for the app you want to govern. Virtual keys are in AI → Credentials & Keys.

02

How a policy fits together

Every request goes through the same checks in the same order. Knowing the order helps you decide where each limit belongs.

The path of one request
  1. 1Identify the caller

    Key, team, app

  2. 2Find the policy

    Most specific binding wins

  3. 3Check the rules

    Models, tools, size, rates

  4. 4Protect the content

    Guardrails

  5. 5Reuse or route

    Cache, then the right provider

  6. 6Reserve budget

    Hold spend before the call

  7. 7Call the provider

A policy supplies the rules for steps two to six. If a request breaks a rule in an enforcing policy, it stops there, and the provider never sees it.

03

Create the policy

Open AI → Policies and create a policy. The form groups settings by purpose, and every field has a short explanation.

  1. 1

    Choose Create policy

    You will find it at the top of the Policies tab.

  2. 2

    Name it for what it governs

    A name like support-app-production tells a reviewer what the policy covers. Names are 3 to 128 characters.

  3. 3

    Set Enforcement mode to Monitor only

    The form starts on Enforce. Monitor only records activity and blocks nothing, which is the safe way to tune.

  4. 4

    Leave Model access preset on All models for now

    You will narrow it in the next section, once you know which models the app uses.

  5. 5

    Save

    The policy exists but is not bound to anything yet. You bind it in the second guide.

AI → Policies → Create policy

Policy basics

2Policy name
support-app-production
3Enforcement mode
EnforceMonitor only
4Model access preset
All modelsClaudeOpenAIGeminiCustom
Name the policy and start in Monitor only mode.

04

Choose what is allowed

Start with the models and providers your teams already use. Allow lists and deny lists accept wildcards, and a deny always wins over an allow.

FieldExampleNotes
Allowed providersopenai, anthropicLeave on All to allow any provider
Allowed modelsgpt-4o-mini, claude-*A wildcard matches a family of models
Denied modelsgpt-4-*Checked first. A deny beats any allow.

05

Add limits that match the app

Limits protect you from the unexpected: an oversized prompt, a long answer, a noisy caller, an agent that loops. Pick starting values from how the app behaves today.

LimitA reasonable startProtects against
Max input tokensTwice your largest normal promptA huge prompt sent by accident
Max output tokensWhat your interface can actually showLong, expensive answers
Request rate per minuteTwice your normal peakA noisy caller or a retry storm
Max tool calls and loop iterationsA small number, such as 10An agent that loops on itself

06

Set a budget

Add a daily budget, a monthly budget, or both. Budget enforcement decides what happens at the limit. The budget is shared by every app, team, and key bound to the policy.

Budget enforcementWhat it doesUse it when
Alert only (never blocks)Tracks spend and never stops trafficYou are learning what normal looks like
Block immediately (fast counter)Stops requests once the budget is reached, using a fast counterYou want a quick hard stop. Needs at least one allowed provider and model.
Precise block (atomic DB check)Checks every request against the budget before it runsYou need exact accounting at the limit

07

Decide how failures behave

Fail mode decides what the gateway does if one of its own checks cannot complete, such as a rate-limit lookup.

Block on error (fail closed)

  • The request is stopped when a check cannot run
  • The safe default for sensitive workloads

Allow on error (fail open)

  • The request continues on the checks that did run
  • A fit for low-risk, high-availability workloads

08

Bind it and send a test request

A policy only applies where it is bound. Bind it to the app's key, or to a team or app. The second guide covers how the gateway chooses between several bindings.

  1. 1

    Add a binding

    Choose the binding type that fits, such as App, and add the environment if the policy is only for production.

  2. 2

    Open the Request Simulator

    It is in the Policies tab. Enter a model and, optionally, the key, team, app, and environment. It reports which policy applies and whether the request would be allowed.

  3. 3

    Send a real request

    Use the virtual key. If the key does not carry a team and app, add them as headers.

bash
curl https://api.cloptima.ai/v1/ai/chat/completions \
  -H "Authorization: Bearer $CLOPTIMA_VIRTUAL_KEY" \
  -H "X-Cloptima-Team: support" \
  -H "X-Cloptima-App: support-assistant" \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "Say hello"}]}'
AI → Policies → Request Simulator
2Model
gpt-4o-mini
2Virtual Key (optional)
No key (policy default)
2Team ID
support
2App ID
support-assistant
2Environment
production

Simulation result

Requestopenai/gpt-4o-mini · Would be allowed
Matched policysupport-app-production · observe
Model checkAllowed
CredentialPlatform managed credits
Evaluate
Try a request before you enforce.

09

Switch to Enforce and break it on purpose

When the simulator and your usage look right, change Mode to Enforce. Then confirm the policy actually blocks what it should.

  1. 1

    Switch Mode to Enforce

    Do this on a policy bound to one app first.

  2. 2

    Send a request with a model you did not allow

    Use a model that is on your deny list or not on the allow list.

  3. 3

    Read the response

    You should see HTTP 403 with a reason that names the rule.

HTTP 403
{
  "error": "Your AI request was blocked because this model is not allowed by the active Cloptima policy.",
  "reason": "model_not_allowed",
  "violations": ["model_not_allowed"]
}

10

If something goes wrong

Most first-policy problems come down to a binding or a mode. This table covers the common ones.

What you seeLikely causeFix
403 policy_not_configuredNo binding matches the callerBind the policy to the key, app, or team
403 model_not_allowedThe model is not on the allow listAllow the model, or use one that is allowed
400 Managed AI requests require Cloptima team and app attributionThe request has no team or appSet them on the key, or send the x-cloptima-team and x-cloptima-app headers
429 request_rate_limit_exceededThe per-minute rate limit was reachedRetry shortly, or raise the limit
The policy seems ignoredIt is in Monitor only mode, or bound to a different scopeCheck Mode and the binding

Put This Guide Into Practice

Cloptima automates the strategies described in this guide.

No credit card required
5-minute setup
Free trial