On this page
01
What you'll set up
Provider scans call an outside service, so they add a small cost and a little time. This guide shows you how to put a ceiling on both, and how to decide what happens when a scan cannot run.
- A per-request cost cap with an action above it
- A latency threshold that skips the scan on slow requests
- A fail mode that matches how sensitive the workload is
- A streaming choice that fits how you deliver answers
You need a guardrail profile with a provider scan, from the previous guide.
02
How a scan decision is made
Before a provider scan runs, the gateway asks four questions in order. A no at any step stops the scan from running.
1Did a local rule already block?
Then skip the scan
2Is the estimated cost under your cap?
If not: skip or block
3Is the request already too slow?
If so: skip
4Run the scan
A skipped scan never blocks a request by itself. The request continues on your local rules. Only the block option on the cost cap blocks.
03
Cap the cost of a scan
In the profile's provider section, use Cost control to choose how the cap behaves.
| Cost control | What it does |
|---|---|
| Off | No cap. Every scan runs. |
| Observe | Records when a scan would exceed the cap. Nothing changes. |
| Enforce | Applies your choice when a scan would exceed the cap. |
- 1
Start in Observe
Set Cost control to Observe. Run for a week so you can see how often the cap would trigger, then move to Enforce.
- 2
Set Max scan cost / request (cents)
Choose a ceiling per request. The estimate uses the provider's list price for the text being scanned.
- 3
Choose When the cap is exceeded
Skip the provider scan lets the request continue. Block request stops it.
- 1Cost control
- Observe
- 2Max scan cost / request (cents)
- 12
- 3When the cap is exceeded
- Skip the provider scan
04
Skip the scan when a request is slow
Provider latency budget (ms) sets a threshold. If a request has already used that much time before the scan would start, the scan is skipped and the request continues.
- Your local rules still apply when the scan is skipped
- Leave it empty to never skip on time
- Pick a value just above your normal request time, so only unusual slowness triggers it
05
Choose what happens when a scan cannot run
If the provider times out, rejects the call, or has no usable credential, the policy's fail mode decides.
Block on error (fail closed)
- The request is stopped
- The safe default for sensitive data
- The response names the guardrail service as unavailable
Allow on error (fail open)
- The request continues on your local rules
- A fit for low-risk internal tools
- Cloptima never adds a check you did not configure
Fail mode lives on the policy, so a customer-facing app can fail closed while an internal prototype fails open. It is in the policy form as Fail mode.
06
What each situation does
This table covers a policy that has a provider scan.
| Situation | Block on error | Allow on error |
|---|---|---|
| Provider times out or errors | Blocked | Continues on local rules |
| Provider credential missing or revoked | Blocked | Continues on local rules |
| Cloptima-managed scan cannot start | Blocked | Continues on local rules |
| Cost above cap, skip | Continues without the scan | Continues without the scan |
| Cost above cap, block | Blocked | Blocked |
| Request already slow | Continues without the scan | Continues without the scan |
07
Protect streamed responses
For streamed responses, Streamed responses decides how the provider scans the output.
| Option | What it does | Best for |
|---|---|---|
| Scan in short batches before sending (can block) | Holds output in small batches and checks each before releasing it | Blocking before content reaches the user |
| Scan once after the stream ends (flags only) | Makes one check after the response and records it | Lower cost, when after-the-fact flagging is enough |
Rules set to Redact or Block hold back a short trailing part of the text while they check it. Observe-only rules hold nothing.
08
A good starting setup
If you are unsure, this order works for most teams.
- 1
Use local rules for credentials and your formats
They are free and instant.
- 2
Add one provider scan for prompts
Choose Prompts only first. Add responses when you need them.
- 3
Put the cost cap in Observe for a week
Learn how often it would trigger.
- 4
Move the cap to Enforce, with skip as the action
Cost stays predictable and traffic keeps flowing.
- 5
Choose Block on error for sensitive policies
Keep Allow on error for low-risk tools.
09
If something goes wrong
Most surprises are a cap or threshold set too low.
| What you see | Likely cause | Fix |
|---|---|---|
| The provider scan never runs | The cost cap is lower than a typical scan, or the latency budget is too small | Raise the cap or the latency budget, or set Cost control to Observe |
| Requests blocked with guardrail_cost_threshold_exceeded | Cost control is Enforce with Block request | Switch the action to Skip the provider scan, or raise the cap |
| Requests blocked with the guardrail service unavailable | The scan could not run and the policy fails closed | Fix the provider or credential, or set the policy to Allow on error |
| Streams cut off mid-response | A rule set to Block matched after content was already sent | Use Redact instead of Block for streamed responses, or Observe while you tune |