Fail Closed or Fail Open? Choosing for AI Safety Checks
Every safety check can fail. Decide in advance whether the request stops or continues, and make that choice per workload.
In this post
Every safety check can fail
A guardrail provider has an outage.
A credential expires. A counter lookup times out. Safety checks are software, and software fails, usually at the least convenient moment.
Picture this
At 2 a.m. the safety service behind your customer-facing assistant stops responding. Customers are mid-conversation. Someone had to decide, long before tonight, whether those requests continue unchecked or stop until the service returns. If nobody decided, the system decided for you.
That decision is the fail mode, and it is worth making on purpose.
The two choices
When a check a policy relies on cannot complete, the gateway can stop the request or let it continue.
Both are legitimate. They protect different things.
Fail closed (the default)
- The request is blocked when a check cannot run
- Data stays protected during an outage
- The cost: an outage becomes a visible interruption
Fail open
- The request continues on the checks that did run
- The app stays available during an outage
- The cost: a request may go unchecked until the service returns
In Cloptima the choice lives on the policy. It applies to checks such as rate counters, spend checks, and provider guardrail scans.
Four questions that decide it
You do not need a long risk review.
Four questions get most workloads to an answer.
- 1
What data touches this workload?
If it can contain customer records, regulated data, or secrets, lean closed.
- 2
What does a short outage cost?
If the app is revenue-critical, ask what other safety layers exist before leaning open.
- 3
Is there another layer?
A guardrail that fails open is easier to accept when your local rules still run and the provider filters content too.
- 4
Who is on the other side?
External customers raise the stakes. Internal tools and prototypes raise them less.
| Workload | Lean toward | Why |
|---|---|---|
| Customer-facing assistant with customer data | Fail closed | Protection matters more than a short interruption |
| Internal developer assistant | Fail open | Availability matters, and local rules still run |
| Batch summaries of public documents | Fail open | Low risk, and easy to retry |
| Regulated or financial workflows | Fail closed | An unchecked request is the bigger risk |
What happens, precisely
Vague promises are not useful in an incident.
This is what each situation does for a policy with a provider guardrail scan.
| Situation | Fail closed | Fail open |
|---|---|---|
| Provider times out or returns an error | Blocked: guardrail service unavailable | Continues on your local rules |
| Provider credential missing or revoked | Blocked | Continues on your local rules |
| Cloptima-managed scan cannot start | Blocked | Continues on your local rules |
| Scan cost above your cap, action: skip | Continues without the scan | Continues without the scan |
| Scan cost above your cap, action: block | Blocked | Blocked |
| Request already slow, scan skipped | Continues without the scan | Continues without the scan |
| A local rule matches | Handled by the rule's action | Handled by the rule's action |
Never a silent substitute
One principle sits under all of this.
When a check cannot run, Cloptima does not quietly swap in a different one, and it does not hide what happened.
- Fail closed blocks and says so: the guardrail service is unavailable and the policy fails closed
- Fail open continues on your own local rules, which have already run
- Cloptima never adds detection you did not configure to cover a gap
You decide what is checked. Nothing else is.
Make failures rarer
The best fail mode is the one you rarely use.
A few settings keep a provider scan from becoming a single point of failure.
1Local rules
No network call
2Cost and latency gate
Your cap, your threshold
3Provider scan
Only if still needed
4Model call
5Response checks
As the answer streams
- Run local rules first. A request they block never pays for a provider scan.
- Set a timeout that fits your latency budget. The default is two seconds (three for Bedrock), and the maximum is ten.
- Cap the cost of a scan per request, and choose to skip it or block above the cap.
- Skip the scan when a request is already slow.
- If you use Cloptima-managed scans, keep your AI credits funded.
Setting it per policy
Because fail mode lives on the policy, one organization can run both postures side by side.
resource "cloptima_llm_gateway_policy" "customer_assistant" {
name = "customer-assistant"
mode = "enforce"
fail_mode = "fail_closed"
}
resource "cloptima_llm_gateway_policy" "internal_prototypes" {
name = "internal-prototypes"
mode = "enforce"
fail_mode = "fail_open"
}A sensible default
If you want a starting point, this one holds up for most teams.
- Fail closed for anything that touches customer, regulated, or financial data
- Fail open for low-risk internal tools
- Write the choice down next to the policy, so the next reader knows it was deliberate
- Revisit it when a workload changes owners or starts handling new data