Pay for the Model You Need, Not the One You Defaulted To

Most teams pick one model and one provider early and never look again. Three small habits change that: ask by model name, route by tier, and fail over automatically. You choose how far to go with each.

Cloptima TeamOctober 5, 2026 9 min read
In this post
  1. 01The default nobody reviewed
  2. 02Three questions worth asking again
  3. 03One model name, several providers
  4. 04Your policy stays in charge
  5. 05A cheaper tier for work that does not need the best
  6. 06What tiering is worth, in numbers
  7. 07Roll it out your way
  8. 08When the provider is down
  9. 09Reliability is a cost decision too
  10. 10Budgets count the retries
  11. 11Roll out in a month
  12. 12Your first week

The default nobody reviewed

Every AI app starts with a choice made in a hurry.

Someone picks the best model they know and the provider they already have an account with. It works, so it ships.

Picture this

A support assistant launches on the largest model from one provider because it gave the best answers in testing. A year later it handles thousands of simple questions a day, such as order status and password resets. Each one still goes to the largest model, at the largest model's price.

Nothing is broken. The choice just stopped being a choice. Nobody asked whether a smaller model would answer the easy questions as well, or whether another provider sells the same model for less.

Three questions worth asking again

Model choice is not one decision.

It is three, and each has its own control.

QuestionWhat answers it
Which provider should serve this model?Ask by model name and let Cloptima choose from the providers you can use
Which model should serve this request?Adaptive routing, with cheap, balanced, and strong tiers you define
What if the provider is down?Retries and automatic fallback to another provider

You decide how far to go with each. They work on their own, and they work together.

One model name, several providers

The same model is often sold by more than one provider.

Claude runs on Anthropic, Vertex AI, and Bedrock. GPT models run on OpenAI and Azure. Naming the provider in your code ties the app to one of them.

Pinned to a provider

  • The code says anthropic/claude-opus-5-5
  • A cheaper route needs a code change
  • An outage at that provider is your outage

Asking by model name

  • The code says claude-opus-5-5
  • Cloptima picks the lowest-cost provider you can use
  • Your policy lists decide what is allowed

The choice is stable. The same kind of request keeps landing on the same provider, which keeps its prompt cache warm, until prices or health change. Your own keys come first, and Cloptima credits cover the rest.

Your policy stays in charge

Choosing a provider for you does not mean choosing around you.

The policy is the authority.

  • Allowed providers keep traffic where your data rules say it belongs
  • Denied models are never chosen
  • A provider you have no credential for is never chosen
  • The Request Simulator shows the choice before you send anything

A cheaper tier for work that does not need the best

Adaptive routing lets you define three tiers of models, cheap, balanced, and strong, and let Cloptima pick a tier per request.

TierA good fitExample work
CheapShort, simple, high-volume requestsOrder status, classification, extraction
BalancedEveryday reasoningSummaries, drafts, routine analysis
StrongHard or high-stakes requestsComplex reasoning, long documents

A model only counts as a candidate if it can do what the request needs: tool calls, structured output, images, streaming, and enough room in its context window. Your policy lists and credentials apply to every candidate.

What tiering is worth, in numbers

The saving depends on how much of your traffic is easy.

Here is an illustration for an assistant that handles 10,000 requests a day.

PlanRequestsCost per requestDaily cost
Everything on the strong model10,000$0.0100$100.00
Cheap tier for simple requests7,000$0.0004$2.80
Balanced tier for everyday requests2,000$0.0030$6.00
Strong tier for hard requests1,000$0.0100$10.00
Tiered total10,000$18.80

In this illustration the daily cost falls by about 81 percent. Your own mix will differ, which is why Observe mode comes first: it shows your numbers before any traffic moves.

Roll it out your way

You choose how fast to move. Each mode is a step you can stop on.

  1. 1

    Observe

    Cloptima recommends a route for each request and changes nothing. You see what it would have chosen.

  2. 2

    Canary

    A share of traffic, which you set, takes the recommended route. The same request always lands in the same group.

  3. 3

    Enforce

    Every eligible request takes the recommended route.

You can also set rollback thresholds on error rate and fallback rate. If a candidate crosses them, routing to that candidate stops without a deploy.

When the provider is down

Providers have bad hours. The question is whether your users notice.

What happens when a call fails
  1. 1Call fails

    Timeout, rate limit, or server error

  2. 2Retry

    With backoff, within your limits

  3. 3Fall back

    To another provider serving the model

  4. 4Answer returned

    No change in your code

When you ask by model name, Cloptima can fall back across the other providers that serve the model, each with its own credential. A fallback never starts after the first words of an answer have been sent, so a reply is never stitched together from two providers.

Reliability is a cost decision too

A failed call costs time, and a retry costs money.

Treat both as settings, not accidents.

SettingTrade-off
More attempts per providerFewer visible failures, more possible cost per request
More fallback providersBetter availability, more price variation if failover happens
Shorter deadlinesFaster failover, more chances of abandoning a slow but good call
Backoff with jitterFewer synchronized retries, a slightly longer tail

Each setting has a default. You change one when you have a reason.

Budgets count the retries

More attempts mean more potential cost.

Budgets account for that.

A budget reserves the worst case for a request, including every attempt your retry settings allow. Reliability never quietly spends past a limit.

Roll out in a month

A staged plan keeps each change small.

  1. 1

    Week 1: ask by model name

    Replace provider-pinned names with plain model names on one app. Check the choice in the Request Simulator.

  2. 2

    Week 2: connect a second provider

    Add a credential, then watch spend by Provider in the Explorer.

  3. 3

    Week 3: observe adaptive routing

    Define your tiers and read what Cloptima would have chosen.

  4. 4

    Week 4: canary

    Send a small share through the recommended route and widen as you gain confidence.

Your first week

Start where the savings and the risk are both small.

  • Pick one app and one model name that several providers serve
  • Set Allowed providers on its policy to match your data rules
  • Add a credential for a second provider
  • Compare spend by Provider and Model in the Explorer after a few days
  • Write down which requests are simple enough for a cheaper tier
Put it into practiceAsk for a model, let Cloptima choose the providerUse one model name across providers, and keep your policy lists in charge.

Keep reading

Ready to Try Cloptima?

Bring LLM FinOps, governed model access, and cloud cost optimization into one operating model.

No credit card required
5-minute setup
Free trial