On this page
- 01What you'll set up
- 02Key scope: who shares an answer
- 03Set the key scope
- 04Content classes: separate kinds of traffic
- 05Sensitive data: bypass or block
- 06Streaming: three choices
- 07Cacheable models and routes
- 08Limits, compression, and timeouts
- 09When the cache starts fresh
- 10A tuning routine
- 11If something goes wrong
01
What you'll set up
A cache that is too cautious saves little. One that is too loose shares what it should not. In about ten minutes you will tune the exact cache for both: who may share entries, what is left out, and how streams and sensitive data are handled.
- Cache key scope that matches how you want to share
- Content classes that separate kinds of traffic
- A choice for requests that contain sensitive data
- A choice for streamed responses
- Limits, compression, and a lookup timeout
You need an exact cache already running, at least in Observe. If not, start with the guide on caching identical requests.
03
Set the key scope
Cache key scope dimensions are on the Exact response cache card.
- 1
Open the card
Go to AI → Policies, open the policy, and find Exact response cache on the Expert step.
- 2
Open Cache key scope dimensions
Leave everything selected to keep entries as narrow as possible.
- 3
Remove what you want to share across
Remove App to share between apps, for example.
- 4
Save and watch
Compare the hit rate in Observe before you enforce the broader scope.
04
Content classes: separate kinds of traffic
A content class is a label you put on a request. Cache only the classes you list.
{
"model": "gpt-4o-mini",
"messages": [{ "role": "user", "content": "What are your opening hours?" }],
"metadata": { "content_class": "faq" }
}- 1
Label your requests
Send metadata.content_class on requests, such as faq, docs_qa, or classification.
- 2
List the classes to cache
In Content classes, enter the labels separated by commas.
- 3
Leave it blank to cache any class
Blank means every class is eligible.
05
Sensitive data: bypass or block
When a request contains data your guardrails flag, you choose what the cache does.
| Choice | What happens |
|---|---|
| Bypass cache | The request goes to the provider as usual and the answer is not stored |
| Block request | The request is refused |
Bypass is the gentle choice and the usual one. Block suits workloads where a request with sensitive data should never be sent at all.
06
Streaming: three choices
Streamed replies reach the user as they are written. Pick how the cache treats them.
| Choice | What happens | Use it when |
|---|---|---|
| Stream and cache after completion | The reply streams normally, and the finished answer is stored | You want repeats of streamed requests to hit |
| Bypass cache | Streams are never cached | Streaming is rare or unique |
| Block request | Streaming requests are refused | The workload must not stream |
07
Cacheable models and routes
By default the cache follows the policy's models. Narrow it when only some models repeat.
- Cacheable models: list models by plain name to admit every provider's offering, or provider/model to pin one
- Routes: the cache covers chat completions, completions, embeddings, Responses, and Messages requests
- Requests that use tools the provider runs for you are never cached
08
Limits, compression, and timeouts
Three settings control cost and speed.
| Setting | Range | Notes |
|---|---|---|
| Max payload bytes | 100 bytes to 10 MB | Larger responses are not stored |
| Compress entries | On or off | Smaller storage for larger responses |
| Compression threshold | 0 to 10 MB | Compress only above this size |
| Lookup timeout | 1 to 5,000 ms | If the lookup is slow, the request goes to the provider |
A slow or unavailable cache never blocks a request. It simply behaves as a miss.
09
When the cache starts fresh
Answers must never outlive the rules that produced them.
- Changing the policy starts the cache fresh for that policy
- Changing a guardrail profile does the same, because the profile is part of the key
- The time to live always applies, so old answers expire on their own
- Invalidate cache, on the policy's row in AI → Policies, clears every stored answer for that policy at once
10
A tuning routine
Change one setting at a time, in Observe.
- 1
Note the hit estimate
Read the preview for the current scope.
- 2
Broaden one thing
Drop App from the scope, or raise the time to live.
- 3
Wait a few days
Compare the estimate.
- 4
Keep what helps
Undo what did not.
11
If something goes wrong
Most surprises come from scope or classes.
| What you see | Likely cause | Fix |
|---|---|---|
| Apps do not share answers | App is still in the key scope | Remove App from Cache key scope dimensions |
| Nothing is cached for a class | The class is not in Content classes | Add it, or leave the list blank |
| Requests with sensitive data fail | Sensitive data is set to Block request | Choose Bypass cache |
| Streamed requests never hit | Streaming is set to Bypass | Choose Stream and cache after completion |