Fail Closed or Fail Open? Choosing for AI Safety Checks

Every safety check can fail. Decide in advance whether the request stops or continues, and make that choice per workload.

Cloptima TeamOctober 5, 2026 7 min read
In this post
  1. 01Every safety check can fail
  2. 02The two choices
  3. 03Four questions that decide it
  4. 04What happens, precisely
  5. 05Never a silent substitute
  6. 06Make failures rarer
  7. 07Setting it per policy
  8. 08A sensible default

Every safety check can fail

A guardrail provider has an outage.

A credential expires. A counter lookup times out. Safety checks are software, and software fails, usually at the least convenient moment.

Picture this

At 2 a.m. the safety service behind your customer-facing assistant stops responding. Customers are mid-conversation. Someone had to decide, long before tonight, whether those requests continue unchecked or stop until the service returns. If nobody decided, the system decided for you.

That decision is the fail mode, and it is worth making on purpose.

The two choices

When a check a policy relies on cannot complete, the gateway can stop the request or let it continue.

Both are legitimate. They protect different things.

Fail closed (the default)

  • The request is blocked when a check cannot run
  • Data stays protected during an outage
  • The cost: an outage becomes a visible interruption

Fail open

  • The request continues on the checks that did run
  • The app stays available during an outage
  • The cost: a request may go unchecked until the service returns

In Cloptima the choice lives on the policy. It applies to checks such as rate counters, spend checks, and provider guardrail scans.

Four questions that decide it

You do not need a long risk review.

Four questions get most workloads to an answer.

  1. 1

    What data touches this workload?

    If it can contain customer records, regulated data, or secrets, lean closed.

  2. 2

    What does a short outage cost?

    If the app is revenue-critical, ask what other safety layers exist before leaning open.

  3. 3

    Is there another layer?

    A guardrail that fails open is easier to accept when your local rules still run and the provider filters content too.

  4. 4

    Who is on the other side?

    External customers raise the stakes. Internal tools and prototypes raise them less.

WorkloadLean towardWhy
Customer-facing assistant with customer dataFail closedProtection matters more than a short interruption
Internal developer assistantFail openAvailability matters, and local rules still run
Batch summaries of public documentsFail openLow risk, and easy to retry
Regulated or financial workflowsFail closedAn unchecked request is the bigger risk
Starting points, not rules

What happens, precisely

Vague promises are not useful in an incident.

This is what each situation does for a policy with a provider guardrail scan.

SituationFail closedFail open
Provider times out or returns an errorBlocked: guardrail service unavailableContinues on your local rules
Provider credential missing or revokedBlockedContinues on your local rules
Cloptima-managed scan cannot startBlockedContinues on your local rules
Scan cost above your cap, action: skipContinues without the scanContinues without the scan
Scan cost above your cap, action: blockBlockedBlocked
Request already slow, scan skippedContinues without the scanContinues without the scan
A local rule matchesHandled by the rule's actionHandled by the rule's action

Never a silent substitute

One principle sits under all of this.

When a check cannot run, Cloptima does not quietly swap in a different one, and it does not hide what happened.

  • Fail closed blocks and says so: the guardrail service is unavailable and the policy fails closed
  • Fail open continues on your own local rules, which have already run
  • Cloptima never adds detection you did not configure to cover a gap

You decide what is checked. Nothing else is.

Make failures rarer

The best fail mode is the one you rarely use.

A few settings keep a provider scan from becoming a single point of failure.

Layered checks
  1. 1Local rules

    No network call

  2. 2Cost and latency gate

    Your cap, your threshold

  3. 3Provider scan

    Only if still needed

  4. 4Model call

  5. 5Response checks

    As the answer streams

  • Run local rules first. A request they block never pays for a provider scan.
  • Set a timeout that fits your latency budget. The default is two seconds (three for Bedrock), and the maximum is ten.
  • Cap the cost of a scan per request, and choose to skip it or block above the cap.
  • Skip the scan when a request is already slow.
  • If you use Cloptima-managed scans, keep your AI credits funded.

Setting it per policy

Because fail mode lives on the policy, one organization can run both postures side by side.

main.tf
resource "cloptima_llm_gateway_policy" "customer_assistant" {
  name      = "customer-assistant"
  mode      = "enforce"
  fail_mode = "fail_closed"
}

resource "cloptima_llm_gateway_policy" "internal_prototypes" {
  name      = "internal-prototypes"
  mode      = "enforce"
  fail_mode = "fail_open"
}

A sensible default

If you want a starting point, this one holds up for most teams.

  • Fail closed for anything that touches customer, regulated, or financial data
  • Fail open for low-risk internal tools
  • Write the choice down next to the policy, so the next reader knows it was deliberate
  • Revisit it when a workload changes owners or starts handling new data
Put it into practiceControl guardrail cost and failure behaviorSet cost caps, skip scans when requests are slow, and choose what happens when a scan cannot run.

Keep reading

Ready to Try Cloptima?

Bring LLM FinOps, governed model access, and cloud cost optimization into one operating model.

No credit card required
5-minute setup
Free trial