All guides

Keep Requests Working When a Provider Fails

Set retries, backoff, deadlines, and fallback so a slow or failing provider becomes a short delay instead of an outage, with budgets that count every attempt.

11 min read Updated October 2026LLM FinOps
On this page
  1. 01What you'll set up
  2. 02Retry, then fall back
  3. 03What counts as a failure
  4. 04Fallback when you ask by model name
  5. 05Set retries, backoff, and deadlines
  6. 06Two deadlines, two jobs
  7. 07Name your own fallback routes
  8. 08Why a reply never switches provider mid-answer
  9. 09Budgets count every attempt
  10. 10Test failover before you need it
  11. 11Roll it out
  12. 12If something goes wrong

01

What you'll set up

Providers have bad hours: timeouts, rate limits, and server errors. In about eleven minutes you will decide how Cloptima retries a failing call, when it moves to another provider, and how long a request may take in total.

  • Retries with backoff for short failures
  • Fallback to another provider for longer ones
  • Deadlines for a whole request and for the first words of an answer
  • Budgets that count every attempt
  • A test that proves failover works before you need it

You need an owner or admin role and a policy bound to your traffic. Every setting here is yours to choose, and each has a sensible default.

02

Retry, then fall back

A failed call goes through two steps. First it is retried on the same provider, with a pause between tries. If that does not work, the request moves to the next provider in line.

What happens when a call fails
  1. 1Call fails

    Timeout, rate limit, server error

  2. 2Retry

    Same provider, with backoff

  3. 3Fall back

    Next provider that serves the model

  4. 4Answer returned

    Your code does not change

03

What counts as a failure

You choose which failures are worth a retry and which should move the request on. The names below are the ones the settings use.

FailureMeaningRetry on the same providerMove to the next provider
connectThe provider could not be reachedYesAfter retries
timeout_before_outputNo output arrived before the deadlineYesAfter retries
rate_limitThe provider is limiting you (HTTP 429)YesYes
provider_5xxThe provider had a server errorYesAfter retries
authenticationThe credential was rejectedYes, with another credential if you have oneYes
quotaThe account hit its quotaYes, with another credential if you have oneYes
model_unavailableThe provider does not serve the model right nowNoYes

Fallback conditions also accept familiar status codes as shorthand: 401, 402, 403, 404, 408, 429, 5xx, and timeout.

04

Fallback when you ask by model name

If your requests name a model without a provider, fallback works with no configuration. Cloptima falls back across the other providers that serve the model.

  • Each fallback provider uses its own credential, or Cloptima credits where they apply
  • A provider with no usable credential is left out of the plan, so it can never be tried without one
  • Your policy lists decide which providers are allowed
  • Where a budget needs a price, a provider with no known price is left out

The guide on asking for a model by name shows how providers are chosen and ordered.

05

Set retries, backoff, and deadlines

Reliability settings sit with the policy's routing settings. Each has a range, and each can be left at its default.

  1. 1

    Open the policy

    Go to AI → Policies and open the policy that governs the traffic.

  2. 2

    Open the retries and fallback settings

    Find the section for retries, backoff, and deadlines.

  3. 3

    Set attempts

    Choose how many tries each provider gets and how many attempts a request may use in total.

  4. 4

    Set backoff

    Choose none, or exponential with jitter, and the shortest and longest pause.

  5. 5

    Set deadlines

    Choose the total time a request may take and how long to wait for the first output.

  6. 6

    Save

    The settings apply to new requests at once.

SettingAllowed rangeWhat it does
Attempts per route1 to 4Tries on one provider before moving on
Total attempts1 to 8The most attempts across all providers for one request
Backoffnone or exponential with jitterHow the pause grows between retries
Initial and maximum backoff0 to 5,000 msThe shortest and longest pause
Request timeoutup to 120 secondsThe deadline for the whole request, across all attempts
First output timeoutup to 60 secondsHow long to wait for the first words before treating the call as failed

06

Two deadlines, two jobs

A request timeout bounds the total wait. A first output timeout catches a call that is stuck before it says anything.

Request timeout

  • Covers every attempt together
  • Keeps a request from retrying forever
  • Set it to the longest wait your users accept

First output timeout

  • Covers one attempt, until output starts
  • Lets a stuck call fail over quickly
  • Set it just above a normal first response

07

Name your own fallback routes

If you want a specific order, list the fallbacks yourself, as provider/model strings. They are tried in order after the request's own route.

main.tf
resource "cloptima_llm_gateway_policy" "support_production" {
  name                           = "support-production"
  mode                           = "enforce"
  allowed_providers              = ["anthropic", "bedrock", "vertex_ai"]
  routing_fallback_routes        = ["bedrock/anthropic.claude-opus-5-5", "vertex_ai/claude-opus-5-5"]
  routing_fallback_conditions    = ["rate_limit", "provider_unavailable", "model_unavailable"]
  routing_max_attempts_per_route = 2
  routing_max_total_attempts     = 5
  routing_backoff                = "exponential_jitter"
  routing_request_timeout_ms     = 30000
  routing_first_output_timeout_ms = 8000
}
  • Up to seven fallback routes
  • Each route needs a credential for its provider
  • The policy is authoritative: a route that the policy's allowed and denied lists exclude is not saved and is never used

08

Why a reply never switches provider mid-answer

Once the first words of an answer have gone to your user, Cloptima does not start a fallback.

Switching providers halfway would stitch one reply from two models. Fallback happens only while nothing has been sent. Streaming replies stay whole, and the first output timeout is what gives a stuck call a chance to fail over before it starts.

09

Budgets count every attempt

More attempts mean more possible cost, and budgets plan for it.

A budget reserves the worst case for a request: the most expensive attempts your settings allow, up to your total attempts. Reliability never quietly spends past a limit.

Setting you raiseEffect on the reserved worst case
Total attemptsReserves room for more attempts
Fallback routesReserves room for the pricier providers you added
Maximum output sizeRaises the cost of each attempt

10

Test failover before you need it

Prove it with a throwaway policy and a credential that fails.

  1. 1

    Create a test policy

    Allow two providers that serve the same model.

  2. 2

    Bind a bad key to the first provider

    Add a credential with a wrong key for the primary provider, and a good one for the second.

  3. 3

    Ask by model name

    Send a request with the plain model name.

  4. 4

    Check the answer and the Explorer

    The request succeeds. Group by Provider and the spend shows on the second provider.

  5. 5

    Clean up

    Fix or remove the bad credential.

11

Roll it out

Reliability settings are safe to add gradually.

  1. 1

    Start with defaults

    Turn on fallback by asking for models by name.

  2. 2

    Add backoff and deadlines

    Set a request timeout your users accept.

  3. 3

    Add explicit fallbacks

    Only where you need a specific order.

  4. 4

    Watch Provider spend

    A sudden shift in the Explorer's Provider grouping means failover is happening.

12

If something goes wrong

Most surprises come from credentials or deadlines.

What you seeLikely causeFix
Requests still fail when a provider is downNo second provider you can useAdd a credential for another provider that serves the model
A fallback route is rejected on saveThe policy's lists exclude itAdd the provider or model to the allowed lists, or choose another route
Requests take too longTotal attempts or timeouts are highLower the request timeout and the total attempts
Spend shifts to a pricier providerFailover is working during an incidentCheck the provider's status, then compare Provider spend
Retries seem to do nothingThe failure type is not in the retry listAdd the failure type to the retry settings

Put This Guide Into Practice

Cloptima automates the strategies described in this guide.

No credit card required
5-minute setup
Free trial