The Cheapest Request Is the One You Do Not Make

Every AI app answers some questions again and again. Caching turns the repeats into free, instant answers. Here is how three layers of caching work, what is safe to cache, and how to prove it pays before you serve anything.

Cloptima TeamOctober 5, 2026 10 min read
In this post
  1. 01The question asked ten thousand times
  2. 02Three layers, three levels of trust
  3. 03Layer one: the prompt prefix
  4. 04Layer two: identical requests
  5. 05Layer three: similar requests
  6. 06What is safe to cache
  7. 07What caching is worth, in numbers
  8. 08What caching costs
  9. 09Prove it pays before you serve anything
  10. 10Measure it on the Dashboard
  11. 11Roll out in a month
  12. 12Your first week

The question asked ten thousand times

Look at what your AI apps are asked on a normal day.

Much of it is not new.

Picture this

A help-center assistant answers 'What are your opening hours?' about four hundred times a day, worded a dozen ways. A classifier labels the same fifty product names thousands of times. Each call is billed and each call takes a second or two, for an answer the system already knew yesterday.

Nobody chose to pay for repeats. Nothing in the app asked whether it had answered this before. A cache is the place that asks.

Three layers, three levels of trust

Caching is not one feature. It is three, and each trades savings against trust differently.

LayerWhat repeatsTrust you extend
Provider prompt cacheThe start of a long promptNone. The provider still writes a fresh answer
Exact response cacheAn identical requestLow. The answer is exactly what you got before
Semantic response cacheA request that means the same thingHigher. You decide how close counts as the same

Start at the top and move down only when the numbers justify it.

Layer one: the prompt prefix

Providers charge less when a long prompt starts the same way as a recent one.

Put fixed instructions and reference text first, and variable text last.

  • No setting to turn on
  • Still a fresh answer every time
  • Visible as savings on the Dashboard

It is the cheapest saving to get. Nothing is reused except the prompt, so nothing can go stale.

Layer two: identical requests

The exact cache stores an answer and returns it when the same request comes again.

Same prompt, same system prompt, same tools, same settings, same model.

Without the exact cache

  • Every repeat is billed
  • Every repeat waits for the provider
  • Spend grows with traffic, not with new questions

With the exact cache

  • A repeat costs nothing
  • A hit is answered in the time of a lookup
  • Spend grows with new questions

A hit carries a response header, so you can see it in a test or a log.

Layer three: similar requests

People rarely repeat a question word for word.

The semantic cache matches by meaning, so 'How do I reset my password' and 'I forgot my password' share one answer.

You choose how close a match has to be, how old a stored answer may be, which apps may share, and which kinds of traffic are eligible. A match must also come from the same model, with the same tools and settings.

What is safe to cache

Safe depends on your content, and you decide.

Cloptima gives you the controls.

ContentGood fit for caching?
Fixed reference answersYes, with a long time to live
Classification of repeated inputsYes
Documentation and FAQ questionsYes, exact first, then semantic
Answers that depend on live dataShort time to live, or no
Requests with personal dataBypass the cache, or block them
  • Entries are kept apart per customer and per credential, and stored encrypted
  • A policy change starts the cache fresh, so old answers cannot outlive the rules behind them
  • Requests with sensitive data can bypass the cache or be refused, as you choose
  • Streamed replies are cached only if you say so, and only once they are complete

What caching is worth, in numbers

The saving depends on how much of your traffic repeats and what each call costs.

Here is an illustration for two apps.

AppMonthly callsCost per callRepeat shareSaved per month
FAQ assistant100,000$0.00455%$220
Classifier50,000$0.00263%$63
Research assistant20,000$0.0812%$192

The largest saving does not belong to the app with the most repeats. It belongs to the app with the costliest calls. That is why you measure in dollars, app by app.

What caching costs

A cache is not free. Knowing the price lets you spend it wisely.

CostWhat to do about it
A stale answerSet a time to live that matches how fast your content changes
A wrong match, in the semantic cacheRaise the similarity threshold and read real matches in Observe first
A small lookup costThe exact cache has none to speak of. The semantic cache uses a small embedding call
Extra things to watchA weekly look at the savings card is enough

None of these is a reason to skip caching. Each one is a setting you control.

Prove it pays before you serve anything

Every cache starts in Observe.

It records what it would have done and serves nothing.

  1. 1

    Observe

    The exact cache shows a savings estimate built from your last 30 days of traffic.

  2. 2

    Enforce on one app

    Compare real savings with the estimate.

  3. 3

    Widen

    Add apps as the numbers hold.

If the estimate is small, you have saved yourself an unneeded cache. If it is large, you know where to start.

Measure it on the Dashboard

Realized Caching Savings splits the savings by layer, and shows a total for the month.

  • Provider prompt cache: what the provider's own caching saved
  • Cloptima exact cache: what answered requests saved
  • Total saved this month

A cache you cannot measure is a cache you cannot defend.

Roll out in a month

A staged plan keeps each change small.

  1. 1

    Week 1: stable prompts

    Put fixed text first so the provider's cache can help.

  2. 2

    Week 2: exact cache in Observe

    Read the preview on your busiest app.

  3. 3

    Week 3: exact cache in Enforce

    Start with one app and compare real savings.

  4. 4

    Week 4: semantic cache in Observe

    Read real matches and set the threshold.

Your first week

Start where repeats are most obvious.

  • Pick one app with repetitive traffic, such as a classifier or an FAQ bot
  • Turn the exact cache on in Observe
  • Set a time to live that matches how fast your content changes
  • Label requests with a content class if you want to cache only some kinds
  • Read the savings card at the end of the week
Put it into practiceCache identical requests and stop paying twiceTurn on the exact response cache in Observe, then serve repeats from cache.

Keep reading

Ready to Try Cloptima?

Bring LLM FinOps, governed model access, and cloud cost optimization into one operating model.

No credit card required
5-minute setup
Free trial