Pay for the Model You Need, Not the One You Defaulted To
Most teams pick one model and one provider early and never look again. Three small habits change that: ask by model name, route by tier, and fail over automatically. You choose how far to go with each.
In this post
- 01The default nobody reviewed
- 02Three questions worth asking again
- 03One model name, several providers
- 04Your policy stays in charge
- 05A cheaper tier for work that does not need the best
- 06What tiering is worth, in numbers
- 07Roll it out your way
- 08When the provider is down
- 09Reliability is a cost decision too
- 10Budgets count the retries
- 11Roll out in a month
- 12Your first week
The default nobody reviewed
Every AI app starts with a choice made in a hurry.
Someone picks the best model they know and the provider they already have an account with. It works, so it ships.
Picture this
A support assistant launches on the largest model from one provider because it gave the best answers in testing. A year later it handles thousands of simple questions a day, such as order status and password resets. Each one still goes to the largest model, at the largest model's price.
Nothing is broken. The choice just stopped being a choice. Nobody asked whether a smaller model would answer the easy questions as well, or whether another provider sells the same model for less.
Three questions worth asking again
Model choice is not one decision.
It is three, and each has its own control.
| Question | What answers it |
|---|---|
| Which provider should serve this model? | Ask by model name and let Cloptima choose from the providers you can use |
| Which model should serve this request? | Adaptive routing, with cheap, balanced, and strong tiers you define |
| What if the provider is down? | Retries and automatic fallback to another provider |
You decide how far to go with each. They work on their own, and they work together.
One model name, several providers
The same model is often sold by more than one provider.
Claude runs on Anthropic, Vertex AI, and Bedrock. GPT models run on OpenAI and Azure. Naming the provider in your code ties the app to one of them.
Pinned to a provider
- The code says anthropic/claude-opus-5-5
- A cheaper route needs a code change
- An outage at that provider is your outage
Asking by model name
- The code says claude-opus-5-5
- Cloptima picks the lowest-cost provider you can use
- Your policy lists decide what is allowed
The choice is stable. The same kind of request keeps landing on the same provider, which keeps its prompt cache warm, until prices or health change. Your own keys come first, and Cloptima credits cover the rest.
Your policy stays in charge
Choosing a provider for you does not mean choosing around you.
The policy is the authority.
- Allowed providers keep traffic where your data rules say it belongs
- Denied models are never chosen
- A provider you have no credential for is never chosen
- The Request Simulator shows the choice before you send anything
A cheaper tier for work that does not need the best
Adaptive routing lets you define three tiers of models, cheap, balanced, and strong, and let Cloptima pick a tier per request.
| Tier | A good fit | Example work |
|---|---|---|
| Cheap | Short, simple, high-volume requests | Order status, classification, extraction |
| Balanced | Everyday reasoning | Summaries, drafts, routine analysis |
| Strong | Hard or high-stakes requests | Complex reasoning, long documents |
A model only counts as a candidate if it can do what the request needs: tool calls, structured output, images, streaming, and enough room in its context window. Your policy lists and credentials apply to every candidate.
What tiering is worth, in numbers
The saving depends on how much of your traffic is easy.
Here is an illustration for an assistant that handles 10,000 requests a day.
| Plan | Requests | Cost per request | Daily cost |
|---|---|---|---|
| Everything on the strong model | 10,000 | $0.0100 | $100.00 |
| Cheap tier for simple requests | 7,000 | $0.0004 | $2.80 |
| Balanced tier for everyday requests | 2,000 | $0.0030 | $6.00 |
| Strong tier for hard requests | 1,000 | $0.0100 | $10.00 |
| Tiered total | 10,000 | $18.80 |
In this illustration the daily cost falls by about 81 percent. Your own mix will differ, which is why Observe mode comes first: it shows your numbers before any traffic moves.
Roll it out your way
You choose how fast to move. Each mode is a step you can stop on.
- 1
Observe
Cloptima recommends a route for each request and changes nothing. You see what it would have chosen.
- 2
Canary
A share of traffic, which you set, takes the recommended route. The same request always lands in the same group.
- 3
Enforce
Every eligible request takes the recommended route.
You can also set rollback thresholds on error rate and fallback rate. If a candidate crosses them, routing to that candidate stops without a deploy.
When the provider is down
Providers have bad hours. The question is whether your users notice.
1Call fails
Timeout, rate limit, or server error
2Retry
With backoff, within your limits
3Fall back
To another provider serving the model
4Answer returned
No change in your code
When you ask by model name, Cloptima can fall back across the other providers that serve the model, each with its own credential. A fallback never starts after the first words of an answer have been sent, so a reply is never stitched together from two providers.
Reliability is a cost decision too
A failed call costs time, and a retry costs money.
Treat both as settings, not accidents.
| Setting | Trade-off |
|---|---|
| More attempts per provider | Fewer visible failures, more possible cost per request |
| More fallback providers | Better availability, more price variation if failover happens |
| Shorter deadlines | Faster failover, more chances of abandoning a slow but good call |
| Backoff with jitter | Fewer synchronized retries, a slightly longer tail |
Each setting has a default. You change one when you have a reason.
Budgets count the retries
More attempts mean more potential cost.
Budgets account for that.
A budget reserves the worst case for a request, including every attempt your retry settings allow. Reliability never quietly spends past a limit.
Roll out in a month
A staged plan keeps each change small.
- 1
Week 1: ask by model name
Replace provider-pinned names with plain model names on one app. Check the choice in the Request Simulator.
- 2
Week 2: connect a second provider
Add a credential, then watch spend by Provider in the Explorer.
- 3
Week 3: observe adaptive routing
Define your tiers and read what Cloptima would have chosen.
- 4
Week 4: canary
Send a small share through the recommended route and widen as you gain confidence.
Your first week
Start where the savings and the risk are both small.
- Pick one app and one model name that several providers serve
- Set Allowed providers on its policy to match your data rules
- Add a credential for a second provider
- Compare spend by Provider and Model in the Explorer after a few days
- Write down which requests are simple enough for a cheaper tier