On this page
- 01What you'll set up
- 02Retry, then fall back
- 03What counts as a failure
- 04Fallback when you ask by model name
- 05Set retries, backoff, and deadlines
- 06Two deadlines, two jobs
- 07Name your own fallback routes
- 08Why a reply never switches provider mid-answer
- 09Budgets count every attempt
- 10Test failover before you need it
- 11Roll it out
- 12If something goes wrong
01
What you'll set up
Providers have bad hours: timeouts, rate limits, and server errors. In about eleven minutes you will decide how Cloptima retries a failing call, when it moves to another provider, and how long a request may take in total.
- Retries with backoff for short failures
- Fallback to another provider for longer ones
- Deadlines for a whole request and for the first words of an answer
- Budgets that count every attempt
- A test that proves failover works before you need it
You need an owner or admin role and a policy bound to your traffic. Every setting here is yours to choose, and each has a sensible default.
02
Retry, then fall back
A failed call goes through two steps. First it is retried on the same provider, with a pause between tries. If that does not work, the request moves to the next provider in line.
1Call fails
Timeout, rate limit, server error
2Retry
Same provider, with backoff
3Fall back
Next provider that serves the model
4Answer returned
Your code does not change
03
What counts as a failure
You choose which failures are worth a retry and which should move the request on. The names below are the ones the settings use.
| Failure | Meaning | Retry on the same provider | Move to the next provider |
|---|---|---|---|
| connect | The provider could not be reached | Yes | After retries |
| timeout_before_output | No output arrived before the deadline | Yes | After retries |
| rate_limit | The provider is limiting you (HTTP 429) | Yes | Yes |
| provider_5xx | The provider had a server error | Yes | After retries |
| authentication | The credential was rejected | Yes, with another credential if you have one | Yes |
| quota | The account hit its quota | Yes, with another credential if you have one | Yes |
| model_unavailable | The provider does not serve the model right now | No | Yes |
Fallback conditions also accept familiar status codes as shorthand: 401, 402, 403, 404, 408, 429, 5xx, and timeout.
04
Fallback when you ask by model name
If your requests name a model without a provider, fallback works with no configuration. Cloptima falls back across the other providers that serve the model.
- Each fallback provider uses its own credential, or Cloptima credits where they apply
- A provider with no usable credential is left out of the plan, so it can never be tried without one
- Your policy lists decide which providers are allowed
- Where a budget needs a price, a provider with no known price is left out
The guide on asking for a model by name shows how providers are chosen and ordered.
05
Set retries, backoff, and deadlines
Reliability settings sit with the policy's routing settings. Each has a range, and each can be left at its default.
- 1
Open the policy
Go to AI → Policies and open the policy that governs the traffic.
- 2
Open the retries and fallback settings
Find the section for retries, backoff, and deadlines.
- 3
Set attempts
Choose how many tries each provider gets and how many attempts a request may use in total.
- 4
Set backoff
Choose none, or exponential with jitter, and the shortest and longest pause.
- 5
Set deadlines
Choose the total time a request may take and how long to wait for the first output.
- 6
Save
The settings apply to new requests at once.
| Setting | Allowed range | What it does |
|---|---|---|
| Attempts per route | 1 to 4 | Tries on one provider before moving on |
| Total attempts | 1 to 8 | The most attempts across all providers for one request |
| Backoff | none or exponential with jitter | How the pause grows between retries |
| Initial and maximum backoff | 0 to 5,000 ms | The shortest and longest pause |
| Request timeout | up to 120 seconds | The deadline for the whole request, across all attempts |
| First output timeout | up to 60 seconds | How long to wait for the first words before treating the call as failed |
06
Two deadlines, two jobs
A request timeout bounds the total wait. A first output timeout catches a call that is stuck before it says anything.
Request timeout
- Covers every attempt together
- Keeps a request from retrying forever
- Set it to the longest wait your users accept
First output timeout
- Covers one attempt, until output starts
- Lets a stuck call fail over quickly
- Set it just above a normal first response
07
Name your own fallback routes
If you want a specific order, list the fallbacks yourself, as provider/model strings. They are tried in order after the request's own route.
resource "cloptima_llm_gateway_policy" "support_production" {
name = "support-production"
mode = "enforce"
allowed_providers = ["anthropic", "bedrock", "vertex_ai"]
routing_fallback_routes = ["bedrock/anthropic.claude-opus-5-5", "vertex_ai/claude-opus-5-5"]
routing_fallback_conditions = ["rate_limit", "provider_unavailable", "model_unavailable"]
routing_max_attempts_per_route = 2
routing_max_total_attempts = 5
routing_backoff = "exponential_jitter"
routing_request_timeout_ms = 30000
routing_first_output_timeout_ms = 8000
}- Up to seven fallback routes
- Each route needs a credential for its provider
- The policy is authoritative: a route that the policy's allowed and denied lists exclude is not saved and is never used
08
Why a reply never switches provider mid-answer
Once the first words of an answer have gone to your user, Cloptima does not start a fallback.
Switching providers halfway would stitch one reply from two models. Fallback happens only while nothing has been sent. Streaming replies stay whole, and the first output timeout is what gives a stuck call a chance to fail over before it starts.
09
Budgets count every attempt
More attempts mean more possible cost, and budgets plan for it.
A budget reserves the worst case for a request: the most expensive attempts your settings allow, up to your total attempts. Reliability never quietly spends past a limit.
| Setting you raise | Effect on the reserved worst case |
|---|---|
| Total attempts | Reserves room for more attempts |
| Fallback routes | Reserves room for the pricier providers you added |
| Maximum output size | Raises the cost of each attempt |
10
Test failover before you need it
Prove it with a throwaway policy and a credential that fails.
- 1
Create a test policy
Allow two providers that serve the same model.
- 2
Bind a bad key to the first provider
Add a credential with a wrong key for the primary provider, and a good one for the second.
- 3
Ask by model name
Send a request with the plain model name.
- 4
Check the answer and the Explorer
The request succeeds. Group by Provider and the spend shows on the second provider.
- 5
Clean up
Fix or remove the bad credential.
11
Roll it out
Reliability settings are safe to add gradually.
- 1
Start with defaults
Turn on fallback by asking for models by name.
- 2
Add backoff and deadlines
Set a request timeout your users accept.
- 3
Add explicit fallbacks
Only where you need a specific order.
- 4
Watch Provider spend
A sudden shift in the Explorer's Provider grouping means failover is happening.
12
If something goes wrong
Most surprises come from credentials or deadlines.
| What you see | Likely cause | Fix |
|---|---|---|
| Requests still fail when a provider is down | No second provider you can use | Add a credential for another provider that serves the model |
| A fallback route is rejected on save | The policy's lists exclude it | Add the provider or model to the allowed lists, or choose another route |
| Requests take too long | Total attempts or timeouts are high | Lower the request timeout and the total attempts |
| Spend shifts to a pricier provider | Failover is working during an incident | Check the provider's status, then compare Provider spend |
| Retries seem to do nothing | The failure type is not in the retry list | Add the failure type to the retry settings |