On this page
01
What you'll set up
People rarely ask the same question the same way. 'How do I reset my password' and 'I forgot my password, what now' need the same answer, and an exact cache treats them as different. The semantic cache matches by meaning. In about eleven minutes you will turn it on, choose how close a match must be, and move from Observe to serving.
- The semantic cache on a policy, in Observe
- A similarity threshold that fits your content
- Limits on age, app, and content class
- Funding for the embeddings that make matching possible
- A rollout from Observe to Enforce
You need an owner or admin role and a policy bound to your traffic.
02
Exact or semantic
Use both. They solve different problems.
Exact cache
- Matches identical requests only
- No extra cost to look up
- Best for fixed prompts and repeated inputs
Semantic cache
- Matches requests that mean the same thing
- A lookup uses a small embedding call
- Best for questions people word differently
03
How a match is judged
A stored answer is reused only if it passes every check. Each one protects you from a bad match.
| Check | Why it matters |
|---|---|
| Same organization | Answers never cross organizations |
| Same app, or an app you allow | An answer written for one app is not reused by another unless you say so |
| Same model and provider | The answer was written by the model you are now calling |
| Same tools and settings | A different tool list or temperature can change the answer |
| Content class you allow | You decide which kinds of traffic may be reused |
| No sensitive data | Neither the new request nor the stored one may hold it |
| Similar enough | The match must meet your similarity threshold |
| Fresh enough | The stored answer must be within your maximum age |
04
Turn it on in Observe
Observe records the matches Cloptima would have made. Nothing is served.
- 1
Open the policy
Go to AI → Policies and open the policy.
- 2
Open the Semantic response cache card
On the Expert step, switch it on.
- 3
Choose Mode: Observe
Nothing changes for your users.
- 4
Set the similarity threshold
Start at 0.90.
- 5
Set the maximum age
Start at 86,400 seconds, one day.
- 6
Save
Matches start being recorded.
- 2Semantic response cache
- Optional add-on for replaying similar answers on approved content classes. Cloptima scopes this to replayable generation routes internally.On
- 3Mode
- Observe
- 4Similarity threshold
- 0.9
- 5Max age (seconds)
- 86400
05
Pick a similarity threshold
The threshold is a score from 0 to 1. Higher means stricter matches.
| Threshold | Effect | Use it for |
|---|---|---|
| 0.97 and above | Only near-identical wording | Content where small differences change the answer |
| 0.92 to 0.96 | Same question, different phrasing | FAQs, help content, documentation questions |
| 0.85 to 0.91 | Loosely similar questions | Only after you have read real matches in Observe |
Tuning with real matches
In Observe, a threshold of 0.90 matches 'How do I reset my password' with 'I forgot my password'. It also matches 'How do I reset my password' with 'How do I change my email'. Raising the threshold to 0.93 keeps the first match and drops the second.
06
Limit what can be reused
Scope the cache to the traffic you trust it with.
| Setting | What it does |
|---|---|
| Max age (seconds) | Stored answers older than this are not reused |
| Allowed apps | Leave blank to keep answers inside one app, or list apps (or *) to share |
| Content classes | The classes of traffic that may be reused, by metadata.content_class. The console starts at *, any class |
| Freshness policy | Strict stops reuse of an answer flagged stale. Best effort allows it, except where redaction rules changed |
07
Add a quality bar if you want one
Required eval score is optional. When you set it, a stored answer is reused only if it scored at least that high in your evals.
Leave it empty to rely on the similarity threshold alone. Set it, for example to 0.95, when you have an eval dataset and want matches to clear a measured quality bar.
08
Fund the embeddings
Matching by meaning needs a small embedding of each request. You choose who pays.
| Option | How it works |
|---|---|
| Cloptima credits | Embeddings are paid from your credits at list price |
| Your own credential | Embeddings use your Vertex AI or local embedding credential |
09
Observe, Suggest, Enforce
Three modes take you from learning to serving.
| Mode | What happens |
|---|---|
| Observe | Matches are recorded. Nothing is served |
| Suggest | Matches that meet your thresholds are recorded as suggestions you can review. Nothing is served |
| Enforce | Matches are served from cache |
You choose when to move. A response served from the semantic cache carries x-cloptima-cache: semantic-hit.
10
Roll it out
Raise your confidence in steps.
- 1
Observe for a week
Read real matches and note the wrong ones.
- 2
Adjust the threshold
Raise it until the wrong matches drop out.
- 3
Enforce for one app and one class
Start with traffic like FAQs.
- 4
Widen
Add apps and classes as the results hold.
11
If something goes wrong
Most questions are about thresholds and scope.
| What you see | Likely cause | Fix |
|---|---|---|
| No matches in Observe | Traffic is unique, or the threshold is high | Lower the threshold slightly, and check the content classes |
| Wrong answers matched | The threshold is too low | Raise the similarity threshold |
| Nothing is eligible | The content class list is empty | Enter a class or * |
| Another app's answers are not reused | Allowed apps is blank | List the apps that may share |
| Matches vanish after a model change | Answers are tied to the model that wrote them | Let the cache refill on the new model |