The Cheapest Request Is the One You Do Not Make
Every AI app answers some questions again and again. Caching turns the repeats into free, instant answers. Here is how three layers of caching work, what is safe to cache, and how to prove it pays before you serve anything.
In this post
- 01The question asked ten thousand times
- 02Three layers, three levels of trust
- 03Layer one: the prompt prefix
- 04Layer two: identical requests
- 05Layer three: similar requests
- 06What is safe to cache
- 07What caching is worth, in numbers
- 08What caching costs
- 09Prove it pays before you serve anything
- 10Measure it on the Dashboard
- 11Roll out in a month
- 12Your first week
The question asked ten thousand times
Look at what your AI apps are asked on a normal day.
Much of it is not new.
Picture this
A help-center assistant answers 'What are your opening hours?' about four hundred times a day, worded a dozen ways. A classifier labels the same fifty product names thousands of times. Each call is billed and each call takes a second or two, for an answer the system already knew yesterday.
Nobody chose to pay for repeats. Nothing in the app asked whether it had answered this before. A cache is the place that asks.
Three layers, three levels of trust
Caching is not one feature. It is three, and each trades savings against trust differently.
| Layer | What repeats | Trust you extend |
|---|---|---|
| Provider prompt cache | The start of a long prompt | None. The provider still writes a fresh answer |
| Exact response cache | An identical request | Low. The answer is exactly what you got before |
| Semantic response cache | A request that means the same thing | Higher. You decide how close counts as the same |
Start at the top and move down only when the numbers justify it.
Layer one: the prompt prefix
Providers charge less when a long prompt starts the same way as a recent one.
Put fixed instructions and reference text first, and variable text last.
- No setting to turn on
- Still a fresh answer every time
- Visible as savings on the Dashboard
It is the cheapest saving to get. Nothing is reused except the prompt, so nothing can go stale.
Layer two: identical requests
The exact cache stores an answer and returns it when the same request comes again.
Same prompt, same system prompt, same tools, same settings, same model.
Without the exact cache
- Every repeat is billed
- Every repeat waits for the provider
- Spend grows with traffic, not with new questions
With the exact cache
- A repeat costs nothing
- A hit is answered in the time of a lookup
- Spend grows with new questions
A hit carries a response header, so you can see it in a test or a log.
Layer three: similar requests
People rarely repeat a question word for word.
The semantic cache matches by meaning, so 'How do I reset my password' and 'I forgot my password' share one answer.
You choose how close a match has to be, how old a stored answer may be, which apps may share, and which kinds of traffic are eligible. A match must also come from the same model, with the same tools and settings.
What is safe to cache
Safe depends on your content, and you decide.
Cloptima gives you the controls.
| Content | Good fit for caching? |
|---|---|
| Fixed reference answers | Yes, with a long time to live |
| Classification of repeated inputs | Yes |
| Documentation and FAQ questions | Yes, exact first, then semantic |
| Answers that depend on live data | Short time to live, or no |
| Requests with personal data | Bypass the cache, or block them |
- Entries are kept apart per customer and per credential, and stored encrypted
- A policy change starts the cache fresh, so old answers cannot outlive the rules behind them
- Requests with sensitive data can bypass the cache or be refused, as you choose
- Streamed replies are cached only if you say so, and only once they are complete
What caching is worth, in numbers
The saving depends on how much of your traffic repeats and what each call costs.
Here is an illustration for two apps.
| App | Monthly calls | Cost per call | Repeat share | Saved per month |
|---|---|---|---|---|
| FAQ assistant | 100,000 | $0.004 | 55% | $220 |
| Classifier | 50,000 | $0.002 | 63% | $63 |
| Research assistant | 20,000 | $0.08 | 12% | $192 |
The largest saving does not belong to the app with the most repeats. It belongs to the app with the costliest calls. That is why you measure in dollars, app by app.
What caching costs
A cache is not free. Knowing the price lets you spend it wisely.
| Cost | What to do about it |
|---|---|
| A stale answer | Set a time to live that matches how fast your content changes |
| A wrong match, in the semantic cache | Raise the similarity threshold and read real matches in Observe first |
| A small lookup cost | The exact cache has none to speak of. The semantic cache uses a small embedding call |
| Extra things to watch | A weekly look at the savings card is enough |
None of these is a reason to skip caching. Each one is a setting you control.
Prove it pays before you serve anything
Every cache starts in Observe.
It records what it would have done and serves nothing.
- 1
Observe
The exact cache shows a savings estimate built from your last 30 days of traffic.
- 2
Enforce on one app
Compare real savings with the estimate.
- 3
Widen
Add apps as the numbers hold.
If the estimate is small, you have saved yourself an unneeded cache. If it is large, you know where to start.
Measure it on the Dashboard
Realized Caching Savings splits the savings by layer, and shows a total for the month.
- Provider prompt cache: what the provider's own caching saved
- Cloptima exact cache: what answered requests saved
- Total saved this month
A cache you cannot measure is a cache you cannot defend.
Roll out in a month
A staged plan keeps each change small.
- 1
Week 1: stable prompts
Put fixed text first so the provider's cache can help.
- 2
Week 2: exact cache in Observe
Read the preview on your busiest app.
- 3
Week 3: exact cache in Enforce
Start with one app and compare real savings.
- 4
Week 4: semantic cache in Observe
Read real matches and set the threshold.
Your first week
Start where repeats are most obvious.
- Pick one app with repetitive traffic, such as a classifier or an FAQ bot
- Turn the exact cache on in Observe
- Set a time to live that matches how fast your content changes
- Label requests with a content class if you want to cache only some kinds
- Read the savings card at the end of the week