All guides

Tune the Exact Cache for Hit Rate and Safety

Share entries where it is safe, keep them apart where it is not, and decide how sensitive data and streams are handled.

10 min read Updated October 2026LLM FinOps
On this page
  1. 01What you'll set up
  2. 02Key scope: who shares an answer
  3. 03Set the key scope
  4. 04Content classes: separate kinds of traffic
  5. 05Sensitive data: bypass or block
  6. 06Streaming: three choices
  7. 07Cacheable models and routes
  8. 08Limits, compression, and timeouts
  9. 09When the cache starts fresh
  10. 10A tuning routine
  11. 11If something goes wrong

01

What you'll set up

A cache that is too cautious saves little. One that is too loose shares what it should not. In about ten minutes you will tune the exact cache for both: who may share entries, what is left out, and how streams and sensitive data are handled.

  • Cache key scope that matches how you want to share
  • Content classes that separate kinds of traffic
  • A choice for requests that contain sensitive data
  • A choice for streamed responses
  • Limits, compression, and a lookup timeout

You need an exact cache already running, at least in Observe. If not, start with the guide on caching identical requests.

02

Key scope: who shares an answer

By default an answer is stored for one app, team, and environment, and reused only by the same. You can narrow the key to share more.

DimensionKept in the keyDrop it to share across
CustomerAlwaysNever. Organizations are always isolated
CredentialAlwaysNever. A different key never sees the answer
Policy versionAlwaysNever. A policy change starts fresh
AppBy defaultDifferent apps
TeamBy defaultDifferent teams
EnvironmentBy defaultTest and production
Model and providerBy defaultDifferent providers of the same model
OrgBy defaultDifferent organizations in one account

Two apps, one FAQ

A help-center bot and an in-product assistant answer the same fixed questions. Dropping App from the key scope lets them share answers, and both see a higher hit rate. Keeping Customer, Credential, and Policy version in the key means nothing leaves your organization or your rules.

03

Set the key scope

Cache key scope dimensions are on the Exact response cache card.

  1. 1

    Open the card

    Go to AI → Policies, open the policy, and find Exact response cache on the Expert step.

  2. 2

    Open Cache key scope dimensions

    Leave everything selected to keep entries as narrow as possible.

  3. 3

    Remove what you want to share across

    Remove App to share between apps, for example.

  4. 4

    Save and watch

    Compare the hit rate in Observe before you enforce the broader scope.

04

Content classes: separate kinds of traffic

A content class is a label you put on a request. Cache only the classes you list.

json
{
  "model": "gpt-4o-mini",
  "messages": [{ "role": "user", "content": "What are your opening hours?" }],
  "metadata": { "content_class": "faq" }
}
  1. 1

    Label your requests

    Send metadata.content_class on requests, such as faq, docs_qa, or classification.

  2. 2

    List the classes to cache

    In Content classes, enter the labels separated by commas.

  3. 3

    Leave it blank to cache any class

    Blank means every class is eligible.

05

Sensitive data: bypass or block

When a request contains data your guardrails flag, you choose what the cache does.

ChoiceWhat happens
Bypass cacheThe request goes to the provider as usual and the answer is not stored
Block requestThe request is refused

Bypass is the gentle choice and the usual one. Block suits workloads where a request with sensitive data should never be sent at all.

06

Streaming: three choices

Streamed replies reach the user as they are written. Pick how the cache treats them.

ChoiceWhat happensUse it when
Stream and cache after completionThe reply streams normally, and the finished answer is storedYou want repeats of streamed requests to hit
Bypass cacheStreams are never cachedStreaming is rare or unique
Block requestStreaming requests are refusedThe workload must not stream

07

Cacheable models and routes

By default the cache follows the policy's models. Narrow it when only some models repeat.

  • Cacheable models: list models by plain name to admit every provider's offering, or provider/model to pin one
  • Routes: the cache covers chat completions, completions, embeddings, Responses, and Messages requests
  • Requests that use tools the provider runs for you are never cached

08

Limits, compression, and timeouts

Three settings control cost and speed.

SettingRangeNotes
Max payload bytes100 bytes to 10 MBLarger responses are not stored
Compress entriesOn or offSmaller storage for larger responses
Compression threshold0 to 10 MBCompress only above this size
Lookup timeout1 to 5,000 msIf the lookup is slow, the request goes to the provider

A slow or unavailable cache never blocks a request. It simply behaves as a miss.

09

When the cache starts fresh

Answers must never outlive the rules that produced them.

  • Changing the policy starts the cache fresh for that policy
  • Changing a guardrail profile does the same, because the profile is part of the key
  • The time to live always applies, so old answers expire on their own
  • Invalidate cache, on the policy's row in AI → Policies, clears every stored answer for that policy at once

10

A tuning routine

Change one setting at a time, in Observe.

  1. 1

    Note the hit estimate

    Read the preview for the current scope.

  2. 2

    Broaden one thing

    Drop App from the scope, or raise the time to live.

  3. 3

    Wait a few days

    Compare the estimate.

  4. 4

    Keep what helps

    Undo what did not.

11

If something goes wrong

Most surprises come from scope or classes.

What you seeLikely causeFix
Apps do not share answersApp is still in the key scopeRemove App from Cache key scope dimensions
Nothing is cached for a classThe class is not in Content classesAdd it, or leave the list blank
Requests with sensitive data failSensitive data is set to Block requestChoose Bypass cache
Streamed requests never hitStreaming is set to BypassChoose Stream and cache after completion

Put This Guide Into Practice

Cloptima automates the strategies described in this guide.

No credit card required
5-minute setup
Free trial