All guides

Control Guardrail Cost and Failure Behavior

Cap what provider safety scans cost, skip them when requests are slow, and choose what happens when a scan cannot run.

9 min read Updated October 2026LLM FinOps
On this page
  1. 01What you'll set up
  2. 02How a scan decision is made
  3. 03Cap the cost of a scan
  4. 04Skip the scan when a request is slow
  5. 05Choose what happens when a scan cannot run
  6. 06What each situation does
  7. 07Protect streamed responses
  8. 08A good starting setup
  9. 09If something goes wrong

01

What you'll set up

Provider scans call an outside service, so they add a small cost and a little time. This guide shows you how to put a ceiling on both, and how to decide what happens when a scan cannot run.

  • A per-request cost cap with an action above it
  • A latency threshold that skips the scan on slow requests
  • A fail mode that matches how sensitive the workload is
  • A streaming choice that fits how you deliver answers

You need a guardrail profile with a provider scan, from the previous guide.

02

How a scan decision is made

Before a provider scan runs, the gateway asks four questions in order. A no at any step stops the scan from running.

Before each provider scan
  1. 1Did a local rule already block?

    Then skip the scan

  2. 2Is the estimated cost under your cap?

    If not: skip or block

  3. 3Is the request already too slow?

    If so: skip

  4. 4Run the scan

A skipped scan never blocks a request by itself. The request continues on your local rules. Only the block option on the cost cap blocks.

03

Cap the cost of a scan

In the profile's provider section, use Cost control to choose how the cap behaves.

Cost controlWhat it does
OffNo cap. Every scan runs.
ObserveRecords when a scan would exceed the cap. Nothing changes.
EnforceApplies your choice when a scan would exceed the cap.
  1. 1

    Start in Observe

    Set Cost control to Observe. Run for a week so you can see how often the cap would trigger, then move to Enforce.

  2. 2

    Set Max scan cost / request (cents)

    Choose a ceiling per request. The estimate uses the provider's list price for the text being scanned.

  3. 3

    Choose When the cap is exceeded

    Skip the provider scan lets the request continue. Block request stops it.

Guardrail profile → Provider scan
1Cost control
Observe
2Max scan cost / request (cents)
12
3When the cap is exceeded
Skip the provider scan
A cap per request, with skip or block above it.

04

Skip the scan when a request is slow

Provider latency budget (ms) sets a threshold. If a request has already used that much time before the scan would start, the scan is skipped and the request continues.

  • Your local rules still apply when the scan is skipped
  • Leave it empty to never skip on time
  • Pick a value just above your normal request time, so only unusual slowness triggers it

05

Choose what happens when a scan cannot run

If the provider times out, rejects the call, or has no usable credential, the policy's fail mode decides.

Block on error (fail closed)

  • The request is stopped
  • The safe default for sensitive data
  • The response names the guardrail service as unavailable

Allow on error (fail open)

  • The request continues on your local rules
  • A fit for low-risk internal tools
  • Cloptima never adds a check you did not configure

Fail mode lives on the policy, so a customer-facing app can fail closed while an internal prototype fails open. It is in the policy form as Fail mode.

06

What each situation does

This table covers a policy that has a provider scan.

SituationBlock on errorAllow on error
Provider times out or errorsBlockedContinues on local rules
Provider credential missing or revokedBlockedContinues on local rules
Cloptima-managed scan cannot startBlockedContinues on local rules
Cost above cap, skipContinues without the scanContinues without the scan
Cost above cap, blockBlockedBlocked
Request already slowContinues without the scanContinues without the scan

07

Protect streamed responses

For streamed responses, Streamed responses decides how the provider scans the output.

OptionWhat it doesBest for
Scan in short batches before sending (can block)Holds output in small batches and checks each before releasing itBlocking before content reaches the user
Scan once after the stream ends (flags only)Makes one check after the response and records itLower cost, when after-the-fact flagging is enough

Rules set to Redact or Block hold back a short trailing part of the text while they check it. Observe-only rules hold nothing.

08

A good starting setup

If you are unsure, this order works for most teams.

  1. 1

    Use local rules for credentials and your formats

    They are free and instant.

  2. 2

    Add one provider scan for prompts

    Choose Prompts only first. Add responses when you need them.

  3. 3

    Put the cost cap in Observe for a week

    Learn how often it would trigger.

  4. 4

    Move the cap to Enforce, with skip as the action

    Cost stays predictable and traffic keeps flowing.

  5. 5

    Choose Block on error for sensitive policies

    Keep Allow on error for low-risk tools.

09

If something goes wrong

Most surprises are a cap or threshold set too low.

What you seeLikely causeFix
The provider scan never runsThe cost cap is lower than a typical scan, or the latency budget is too smallRaise the cap or the latency budget, or set Cost control to Observe
Requests blocked with guardrail_cost_threshold_exceededCost control is Enforce with Block requestSwitch the action to Skip the provider scan, or raise the cap
Requests blocked with the guardrail service unavailableThe scan could not run and the policy fails closedFix the provider or credential, or set the policy to Allow on error
Streams cut off mid-responseA rule set to Block matched after content was already sentUse Redact instead of Block for streamed responses, or Observe while you tune

Put This Guide Into Practice

Cloptima automates the strategies described in this guide.

No credit card required
5-minute setup
Free trial