On this page
01
What you'll set up
In about ten minutes you will have a policy that limits which models an app can use, how large its requests can be, and how much it can spend per day. You will test it before any live traffic is affected.
- A policy in Monitor only mode, so nothing is blocked while you tune it
- Model access, size limits, and a rate limit that fit your app
- A daily budget with the enforcement style you choose
- A tested path from a simulated request to a real one
You will need an owner or admin role in your organization, an AI plan (policy enforcement is part of the AI plans), and a virtual key for the app you want to govern. Virtual keys are in AI → Credentials & Keys.
02
How a policy fits together
Every request goes through the same checks in the same order. Knowing the order helps you decide where each limit belongs.
1Identify the caller
Key, team, app
2Find the policy
Most specific binding wins
3Check the rules
Models, tools, size, rates
4Protect the content
Guardrails
5Reuse or route
Cache, then the right provider
6Reserve budget
Hold spend before the call
7Call the provider
A policy supplies the rules for steps two to six. If a request breaks a rule in an enforcing policy, it stops there, and the provider never sees it.
03
Create the policy
Open AI → Policies and create a policy. The form groups settings by purpose, and every field has a short explanation.
- 1
Choose Create policy
You will find it at the top of the Policies tab.
- 2
Name it for what it governs
A name like support-app-production tells a reviewer what the policy covers. Names are 3 to 128 characters.
- 3
Set Enforcement mode to Monitor only
The form starts on Enforce. Monitor only records activity and blocks nothing, which is the safe way to tune.
- 4
Leave Model access preset on All models for now
You will narrow it in the next section, once you know which models the app uses.
- 5
Save
The policy exists but is not bound to anything yet. You bind it in the second guide.
Policy basics
- 2Policy name
- support-app-production
- 3Enforcement mode
- EnforceMonitor only
- 4Model access preset
- All modelsClaudeOpenAIGeminiCustom
04
Choose what is allowed
Start with the models and providers your teams already use. Allow lists and deny lists accept wildcards, and a deny always wins over an allow.
| Field | Example | Notes |
|---|---|---|
| Allowed providers | openai, anthropic | Leave on All to allow any provider |
| Allowed models | gpt-4o-mini, claude-* | A wildcard matches a family of models |
| Denied models | gpt-4-* | Checked first. A deny beats any allow. |
05
Add limits that match the app
Limits protect you from the unexpected: an oversized prompt, a long answer, a noisy caller, an agent that loops. Pick starting values from how the app behaves today.
| Limit | A reasonable start | Protects against |
|---|---|---|
| Max input tokens | Twice your largest normal prompt | A huge prompt sent by accident |
| Max output tokens | What your interface can actually show | Long, expensive answers |
| Request rate per minute | Twice your normal peak | A noisy caller or a retry storm |
| Max tool calls and loop iterations | A small number, such as 10 | An agent that loops on itself |
06
Set a budget
Add a daily budget, a monthly budget, or both. Budget enforcement decides what happens at the limit. The budget is shared by every app, team, and key bound to the policy.
| Budget enforcement | What it does | Use it when |
|---|---|---|
| Alert only (never blocks) | Tracks spend and never stops traffic | You are learning what normal looks like |
| Block immediately (fast counter) | Stops requests once the budget is reached, using a fast counter | You want a quick hard stop. Needs at least one allowed provider and model. |
| Precise block (atomic DB check) | Checks every request against the budget before it runs | You need exact accounting at the limit |
07
Decide how failures behave
Fail mode decides what the gateway does if one of its own checks cannot complete, such as a rate-limit lookup.
Block on error (fail closed)
- The request is stopped when a check cannot run
- The safe default for sensitive workloads
Allow on error (fail open)
- The request continues on the checks that did run
- A fit for low-risk, high-availability workloads
08
Bind it and send a test request
A policy only applies where it is bound. Bind it to the app's key, or to a team or app. The second guide covers how the gateway chooses between several bindings.
- 1
Add a binding
Choose the binding type that fits, such as App, and add the environment if the policy is only for production.
- 2
Open the Request Simulator
It is in the Policies tab. Enter a model and, optionally, the key, team, app, and environment. It reports which policy applies and whether the request would be allowed.
- 3
Send a real request
Use the virtual key. If the key does not carry a team and app, add them as headers.
curl https://api.cloptima.ai/v1/ai/chat/completions \
-H "Authorization: Bearer $CLOPTIMA_VIRTUAL_KEY" \
-H "X-Cloptima-Team: support" \
-H "X-Cloptima-App: support-assistant" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "Say hello"}]}'- 2Model
- gpt-4o-mini
- 2Virtual Key (optional)
- No key (policy default)
- 2Team ID
- support
- 2App ID
- support-assistant
- 2Environment
- production
Simulation result
| Request | openai/gpt-4o-mini · Would be allowed |
| Matched policy | support-app-production · observe |
| Model check | Allowed |
| Credential | Platform managed credits |
09
Switch to Enforce and break it on purpose
When the simulator and your usage look right, change Mode to Enforce. Then confirm the policy actually blocks what it should.
- 1
Switch Mode to Enforce
Do this on a policy bound to one app first.
- 2
Send a request with a model you did not allow
Use a model that is on your deny list or not on the allow list.
- 3
Read the response
You should see HTTP 403 with a reason that names the rule.
{
"error": "Your AI request was blocked because this model is not allowed by the active Cloptima policy.",
"reason": "model_not_allowed",
"violations": ["model_not_allowed"]
}10
If something goes wrong
Most first-policy problems come down to a binding or a mode. This table covers the common ones.
| What you see | Likely cause | Fix |
|---|---|---|
| 403 policy_not_configured | No binding matches the caller | Bind the policy to the key, app, or team |
| 403 model_not_allowed | The model is not on the allow list | Allow the model, or use one that is allowed |
| 400 Managed AI requests require Cloptima team and app attribution | The request has no team or app | Set them on the key, or send the x-cloptima-team and x-cloptima-app headers |
| 429 request_rate_limit_exceeded | The per-minute rate limit was reached | Retry shortly, or raise the limit |
| The policy seems ignored | It is in Monitor only mode, or bound to a different scope | Check Mode and the binding |