All guides

Measure What AI Caching Saves

Read the savings from provider caching and the Cloptima cache, compare would-hits with real hits, and find the next app worth caching.

8 min read Updated October 2026LLM FinOps
On this page
  1. 01What you'll do
  2. 02Three layers of savings
  3. 03Read the savings card
  4. 04A worked example
  5. 05Why a hit rate alone can mislead
  6. 06Observe against Enforce
  7. 07Confirm hits in your own tests
  8. 08Find the next app to cache
  9. 09A monthly review
  10. 10If something goes wrong

01

What you'll do

Caching only matters if it pays. In about eight minutes you will read the savings, tell real hits from would-hits, and decide where to cache next.

  • Read the Realized Caching Savings card
  • Tell the three layers of savings apart
  • Compare Observe estimates with Enforce results
  • Find the apps worth caching next

02

Three layers of savings

Three different mechanisms can lower a cost. Keep them apart when you measure.

LayerHow it savesWhere it shows
Provider prompt cacheThe provider charges less for a repeated prompt prefixProvider prompt cache on the savings card
Cloptima exact cacheAn identical request is answered without calling the providerCloptima exact cache on the savings card
Semantic cacheA similar request is answered without calling the providerSemantic cache on the savings card

03

Read the savings card

Open AI → Dashboard and find Realized Caching Savings.

LineMeaning
Provider prompt cacheSavings from your provider's own prompt caching this month
Cloptima exact cacheSavings from requests answered from the exact cache this month
Semantic cache (observed)What semantic matching has found, and its savings once you serve from it
Total saved this monthThe layers added together
Cache read tokens and Cache write tokensHow much prompt text was read from or written to the provider's cache
AI → Dashboard → Realized Caching Savings
Provider prompt cache$128.40
Cloptima exact cache$220.00
Semantic cache (observed)Observe mode — savings not yet quantified
Total saved this month$348.40
Cache read tokens41.2M
Cache write tokens3.8M
Savings by layer, and the total.

04

A worked example

Take a classifier that makes 50,000 calls a month at $0.002 each. The numbers below are an illustration.

MeasureObserve estimateEnforce, first month
Calls that repeat an earlier request34,000 (68%)31,500 (63%)
Cost of those calls at list price$68.00$63.00
Spend without caching$100.00$100.00
Spend with caching$37.00
Saved this month$63.00

Real savings landed a little under the estimate, which is normal: the first hours of a cache are empty, and some repeats come after an entry has expired. A gap of this size is healthy. A gap of half or more means the time to live or key scope is narrower than the estimate assumed.

05

Why a hit rate alone can mislead

A high hit rate on cheap calls saves less than a modest hit rate on costly ones.

AppHit rateCost per callMonthly callsSaved
Classifier63%$0.00250,000$63
Research assistant12%$0.0820,000$192

Read savings in dollars, then ask where a higher hit rate would be worth the most. The research assistant has the lower hit rate and the bigger saving.

06

Observe against Enforce

Observe estimates what a cache would save. Enforce is the real number. Compare them.

  1. 1

    Note the estimate in Observe

    Use the blast-radius preview on the exact cache card.

  2. 2

    Switch to Enforce on one app

    Wait a week of ordinary traffic.

  3. 3

    Compare

    If real savings are well below the estimate, check the time to live and the key scope.

07

Confirm hits in your own tests

The response header lets you check a hit without a dashboard.

Header valueMeaning
x-cloptima-cache: hitServed from the exact cache
x-cloptima-cache: semantic-hitServed from the semantic cache
No headerThe request went to the provider

Send the same request twice. The second should carry the header.

08

Find the next app to cache

Use the Explorer to find where repeated spend lives.

  1. 1

    Group by App

    In AI → Explorer, choose a 30d window and sort by Spend.

  2. 2

    Pick apps with repetitive work

    Classifiers, FAQ bots, and fixed-prompt pipelines repeat the most.

  3. 3

    Run Observe on each

    Read its preview before you enforce.

09

A monthly review

A short review keeps caching honest.

  • Read Total saved this month and compare with last month
  • Check each cache's time to live against how fast your content changes
  • Look at apps that cost a lot and cache nothing
  • Retire a cache that saves little and adds risk

10

If something goes wrong

Most surprises come from mode and scope.

What you seeLikely causeFix
Cloptima exact cache shows zeroThe cache is in Observe, or no requests repeatedCheck the mode, and the preview
Savings are lower than the estimateTime to live or key scope is narrower than the estimate assumedWiden them in Observe and compare
Provider prompt cache shows nothingPrompts do not repeat a stable prefixPut fixed instructions first and variable text last
Semantic cache shows no savingsIt is in Observe or SuggestSwitch to Enforce when you are ready

Put This Guide Into Practice

Cloptima automates the strategies described in this guide.

No credit card required
5-minute setup
Free trial