Why Every AI Request Needs a Policy
AI spend and risk grow where traffic is ungoverned. A policy on every request, with no policy meaning no traffic, is the simplest control that works.
In this post
The AI bill nobody approved
Most AI cost and risk stories begin the same way.
Nobody did anything reckless. A team shipped something useful, and the traffic behind it never had an owner.
Picture this
A support team wires an internal assistant to a premium model on Thursday. It works well, so a second team copies the API key into their own service. On Friday evening a tool-calling agent hits an error and retries itself, calling the model thousands of times before anyone is awake. On Monday there is a bill, a Slack thread, and no clear answer to who was allowed to do what.
None of those steps was a mistake on its own. The gap is that the layer between the apps and the model providers treated every call as acceptable. Three patterns show up again and again:
- The shared key: one credential used by several apps, so spend and incidents cannot be traced to an owner.
- The unapproved model: a team picks the newest, most expensive model because nothing says otherwise.
- The runaway loop: an agent retries or loops with no ceiling on tool calls, retries, or tokens.
Each of these is cheap to prevent before it happens and expensive to untangle afterwards. A policy is how you prevent them.
What a policy is, in one sentence
A policy is a set of rules the gateway checks on every request before it reaches a model provider.
For each request it answers one question: is this caller allowed to do this, at this size, at this cost?
| Control | What you set | What the developer sees when it is broken |
|---|---|---|
| Models and providers | Allow and deny lists, with wildcards such as gpt-* | model_not_allowed, with a plain-language message |
| Tools | Which tools and tool servers an agent may call | tool_not_allowed |
| Size | Input tokens, output tokens, tool calls, retries, loop depth | max_input_tokens_exceeded and similar |
| Rate | Requests and tokens per minute | HTTP 429: retry shortly |
| Spend | Daily and monthly budgets, alert-only or blocking | HTTP 402 when a blocking budget is reached |
| Content | A guardrail profile for secrets and personal data | gateway_guardrail_blocked |
All of it lives in one place, so a security reviewer, a platform engineer, and a finance partner can read the same page and agree on what is allowed.
1Identify the caller
Key, team, app
2Find the policy
Most specific binding wins
3Check the rules
Models, tools, size, rates
4Protect the content
Guardrails on prompts and responses
5Reuse or route
Cache, then the right provider
6Reserve budget
Hold spend before the call
7Call the provider
The response streams back
No policy, no traffic
The most important rule in the gateway is also the simplest.
A request needs a policy. If no policy is bound to the key, app, or team that sent it, the gateway refuses the request and says what to fix.
{
"error": "Your AI request was blocked because no Cloptima policy is bound to this key, app, or team. Bind a policy to allow traffic.",
"reason": "policy_not_configured",
"violations": ["policy_not_configured"]
}There is no hidden default policy to forget about. A permissive default is easy to start with, but it quietly becomes the policy nobody reviews. Starting from a refusal turns shadow AI into a visible error that a developer fixes in a minute.
The safest default is a refusal that tells you exactly what to fix.
Bind rules to the people who own the traffic
A policy only matters where it is bound.
Bind it to a team, an app, an environment, a caller type, or a single key, and leave any field empty to match everything for that field.
| Scope | Policy | Why |
|---|---|---|
| Environment: production | Company production policy | One floor for everything that runs in production |
| Team: research, environment: sandbox | Relaxed policy with a low daily budget | Room to experiment and a hard stop on spend |
| Key: billing-agent | Tight tool allow list and a small output cap | One integration, one set of limits |
When more than one binding matches, the gateway picks the lowest priority number first, then the most specific scope, then the newest binding. If you never change priority, the most specific binding wins. When a new binding overlaps another one and either uses a custom priority, the console shows which would win and asks you to acknowledge it. The acknowledgement lands in the audit log.
Strictest wins, so central rules hold
Policies stack. When several enforcing policies match a request, the gateway applies the strictest request size, rate, tool, and budget limit across all of them, and honors every deny list.
A concrete case
Security creates an enforcing company-wide policy that denies a shell-execution tool. The data team has its own policy that allows it. A request from the data team matches both. The tool is still denied, because a deny from any enforcing policy applies.
Without strictest-wins
- A team can loosen a company rule by writing its own policy
- Security has to audit every team policy
- The most specific policy overrides the floor
With strictest-wins
- Security sets the floor once
- Teams add rules on top
- A looser team policy cannot remove the floor
Roll out without breaking anyone
You do not need a big-bang change.
The safest rollout touches one app at a time and takes about a week.
- 1
Create the policy in Monitor only mode
Start with the models and providers your teams already use. The gateway records decisions and blocks nothing.
- 2
Try risky requests in the Request Simulator
Enter a model you plan to deny and read the result without sending live traffic.
- 3
Watch real usage
Check spend, models, and usage by app in the console to confirm your limits are realistic.
- 4
Enforce for one app
Switch Mode to Enforce on a policy bound to a single app. A block now returns a clear message.
- 5
Widen the binding
Bind the policy to more apps and teams as confidence grows.
What developers see when something is blocked
A policy is only as good as the error it produces.
A block should tell the developer what happened and how to fix it, not send them to a support queue.
{
"error": "Your AI request was blocked because this model is not allowed by the active Cloptima policy.",
"reason": "model_not_allowed",
"violations": ["model_not_allowed"]
}| Reason | What it means | The fix |
|---|---|---|
| policy_not_configured | No policy is bound to the caller | Bind a policy to the key, app, or team |
| model_not_allowed | The model is not on the allow list | Use an allowed model, or update the list |
| max_input_tokens_exceeded | The request is larger than the policy allows | Send less, or raise the limit |
| request_rate_limit_exceeded | Too many requests this minute (HTTP 429) | Retry shortly, or raise the limit |
Every block also appears in the Policy Violations card in the Audit tab, so security teams can see patterns without asking developers.
Policies as code
Platform teams that keep infrastructure in Terraform can keep policies there too.
A policy, its binding, and the virtual key an app uses can be reviewed in a pull request like anything else.
resource "cloptima_llm_gateway_policy" "production" {
name = "production-default-policy"
mode = "enforce"
allowed_models = ["openai/gpt-4o", "openai/gpt-4o-mini"]
request_rate_limit_per_minute = 120
daily_budget_usd = 500
monthly_budget_usd = 10000
}
resource "cloptima_llm_gateway_policy_binding" "production" {
policy_id = cloptima_llm_gateway_policy.production.id
environment = "production"
}Your first week
You do not need to model your whole organization on day one.
A small, honest first policy beats a perfect one that never ships.
- Create one policy for production with the models you actually use
- Add a maximum input size, a maximum output size, and a request rate
- Add a daily budget so a runaway loop stops the same day
- Bind it to your busiest app and test it in the Request Simulator
- Switch to Enforce once the results look right, then widen the binding