Why Every AI Request Needs a Policy

AI spend and risk grow where traffic is ungoverned. A policy on every request, with no policy meaning no traffic, is the simplest control that works.

Cloptima TeamOctober 5, 2026 8 min read
In this post
  1. 01The AI bill nobody approved
  2. 02What a policy is, in one sentence
  3. 03No policy, no traffic
  4. 04Bind rules to the people who own the traffic
  5. 05Strictest wins, so central rules hold
  6. 06Roll out without breaking anyone
  7. 07What developers see when something is blocked
  8. 08Policies as code
  9. 09Your first week

The AI bill nobody approved

Most AI cost and risk stories begin the same way.

Nobody did anything reckless. A team shipped something useful, and the traffic behind it never had an owner.

Picture this

A support team wires an internal assistant to a premium model on Thursday. It works well, so a second team copies the API key into their own service. On Friday evening a tool-calling agent hits an error and retries itself, calling the model thousands of times before anyone is awake. On Monday there is a bill, a Slack thread, and no clear answer to who was allowed to do what.

None of those steps was a mistake on its own. The gap is that the layer between the apps and the model providers treated every call as acceptable. Three patterns show up again and again:

  • The shared key: one credential used by several apps, so spend and incidents cannot be traced to an owner.
  • The unapproved model: a team picks the newest, most expensive model because nothing says otherwise.
  • The runaway loop: an agent retries or loops with no ceiling on tool calls, retries, or tokens.

Each of these is cheap to prevent before it happens and expensive to untangle afterwards. A policy is how you prevent them.

What a policy is, in one sentence

A policy is a set of rules the gateway checks on every request before it reaches a model provider.

For each request it answers one question: is this caller allowed to do this, at this size, at this cost?

ControlWhat you setWhat the developer sees when it is broken
Models and providersAllow and deny lists, with wildcards such as gpt-*model_not_allowed, with a plain-language message
ToolsWhich tools and tool servers an agent may calltool_not_allowed
SizeInput tokens, output tokens, tool calls, retries, loop depthmax_input_tokens_exceeded and similar
RateRequests and tokens per minuteHTTP 429: retry shortly
SpendDaily and monthly budgets, alert-only or blockingHTTP 402 when a blocking budget is reached
ContentA guardrail profile for secrets and personal datagateway_guardrail_blocked
What a policy can control

All of it lives in one place, so a security reviewer, a platform engineer, and a finance partner can read the same page and agree on what is allowed.

The path of one request
  1. 1Identify the caller

    Key, team, app

  2. 2Find the policy

    Most specific binding wins

  3. 3Check the rules

    Models, tools, size, rates

  4. 4Protect the content

    Guardrails on prompts and responses

  5. 5Reuse or route

    Cache, then the right provider

  6. 6Reserve budget

    Hold spend before the call

  7. 7Call the provider

    The response streams back

No policy, no traffic

The most important rule in the gateway is also the simplest.

A request needs a policy. If no policy is bound to the key, app, or team that sent it, the gateway refuses the request and says what to fix.

HTTP 403
{
  "error": "Your AI request was blocked because no Cloptima policy is bound to this key, app, or team. Bind a policy to allow traffic.",
  "reason": "policy_not_configured",
  "violations": ["policy_not_configured"]
}

There is no hidden default policy to forget about. A permissive default is easy to start with, but it quietly becomes the policy nobody reviews. Starting from a refusal turns shadow AI into a visible error that a developer fixes in a minute.

The safest default is a refusal that tells you exactly what to fix.

Bind rules to the people who own the traffic

A policy only matters where it is bound.

Bind it to a team, an app, an environment, a caller type, or a single key, and leave any field empty to match everything for that field.

ScopePolicyWhy
Environment: productionCompany production policyOne floor for everything that runs in production
Team: research, environment: sandboxRelaxed policy with a low daily budgetRoom to experiment and a hard stop on spend
Key: billing-agentTight tool allow list and a small output capOne integration, one set of limits
Example bindings. The console binds by app, team, or key with an optional environment; environment-only bindings are set with Terraform or the API.

When more than one binding matches, the gateway picks the lowest priority number first, then the most specific scope, then the newest binding. If you never change priority, the most specific binding wins. When a new binding overlaps another one and either uses a custom priority, the console shows which would win and asks you to acknowledge it. The acknowledgement lands in the audit log.

Strictest wins, so central rules hold

Policies stack. When several enforcing policies match a request, the gateway applies the strictest request size, rate, tool, and budget limit across all of them, and honors every deny list.

A concrete case

Security creates an enforcing company-wide policy that denies a shell-execution tool. The data team has its own policy that allows it. A request from the data team matches both. The tool is still denied, because a deny from any enforcing policy applies.

Without strictest-wins

  • A team can loosen a company rule by writing its own policy
  • Security has to audit every team policy
  • The most specific policy overrides the floor

With strictest-wins

  • Security sets the floor once
  • Teams add rules on top
  • A looser team policy cannot remove the floor

Roll out without breaking anyone

You do not need a big-bang change.

The safest rollout touches one app at a time and takes about a week.

  1. 1

    Create the policy in Monitor only mode

    Start with the models and providers your teams already use. The gateway records decisions and blocks nothing.

  2. 2

    Try risky requests in the Request Simulator

    Enter a model you plan to deny and read the result without sending live traffic.

  3. 3

    Watch real usage

    Check spend, models, and usage by app in the console to confirm your limits are realistic.

  4. 4

    Enforce for one app

    Switch Mode to Enforce on a policy bound to a single app. A block now returns a clear message.

  5. 5

    Widen the binding

    Bind the policy to more apps and teams as confidence grows.

What developers see when something is blocked

A policy is only as good as the error it produces.

A block should tell the developer what happened and how to fix it, not send them to a support queue.

HTTP 403
{
  "error": "Your AI request was blocked because this model is not allowed by the active Cloptima policy.",
  "reason": "model_not_allowed",
  "violations": ["model_not_allowed"]
}
ReasonWhat it meansThe fix
policy_not_configuredNo policy is bound to the callerBind a policy to the key, app, or team
model_not_allowedThe model is not on the allow listUse an allowed model, or update the list
max_input_tokens_exceededThe request is larger than the policy allowsSend less, or raise the limit
request_rate_limit_exceededToo many requests this minute (HTTP 429)Retry shortly, or raise the limit

Every block also appears in the Policy Violations card in the Audit tab, so security teams can see patterns without asking developers.

Policies as code

Platform teams that keep infrastructure in Terraform can keep policies there too.

A policy, its binding, and the virtual key an app uses can be reviewed in a pull request like anything else.

main.tf
resource "cloptima_llm_gateway_policy" "production" {
  name                          = "production-default-policy"
  mode                          = "enforce"
  allowed_models                = ["openai/gpt-4o", "openai/gpt-4o-mini"]
  request_rate_limit_per_minute = 120
  daily_budget_usd              = 500
  monthly_budget_usd            = 10000
}

resource "cloptima_llm_gateway_policy_binding" "production" {
  policy_id   = cloptima_llm_gateway_policy.production.id
  environment = "production"
}

Your first week

You do not need to model your whole organization on day one.

A small, honest first policy beats a perfect one that never ships.

  • Create one policy for production with the models you actually use
  • Add a maximum input size, a maximum output size, and a request rate
  • Add a daily budget so a runaway loop stops the same day
  • Bind it to your busiest app and test it in the Request Simulator
  • Switch to Enforce once the results look right, then widen the binding
Put it into practiceCreate your first AI gateway policySet model access, limits, and spend controls in Monitor only mode, then switch to Enforce.

Keep reading

Ready to Try Cloptima?

Bring LLM FinOps, governed model access, and cloud cost optimization into one operating model.

No credit card required
5-minute setup
Free trial