On this page
01
What you'll set up
Many apps send the same request over and over: a standard prompt, a classification of a repeated input, an embedding of a known text. Each repeat costs money and time for an answer you already have. In about eleven minutes you will turn on the exact response cache, watch what it would save, and then start serving repeats from cache.
- The exact response cache on a policy, in Observe
- A savings estimate from your own traffic
- A time to live and size limit that suit your content
- Enforce mode, so repeated requests are answered from cache
- A way to see a hit in the response
You need an owner or admin role and a policy bound to your traffic.
02
How the exact cache works
Before a request goes to a provider, Cloptima checks whether it has answered an identical request before under the same rules. If it has, the stored answer comes back and the provider is never called.
1Request arrives
Policy and guardrails applied
2Cache looked up
Same request, same rules?
3Hit
Stored answer returned, no provider cost
4Miss
Provider answers, the answer is stored
Exact means exact. A request must match on the prompt, the system prompt, the tools, and every setting that changes the answer, such as temperature. Cloptima does not guess that two different requests are close enough.
03
Start in Observe
Observe records every request that would have been a hit and serves nothing. You learn the value of the cache before it changes a single response.
- 1
Open the policy
Go to AI → Policies and open the policy that governs the traffic.
- 2
Open the Exact response cache card
On the Expert step, switch on Exact response cache.
- 3
Choose Mode: Monitor only
This is Observe for the cache. Nothing is served from cache yet.
- 4
Leave the defaults
The time to live, size limit, and lookup timeout have sensible starting values.
- 5
Save
Cloptima starts recording would-hits at once.
- 2Exact response cache
- Optional add-on for fast replay of identical requests. Cloptima handles internal route scoping.On
- 3Mode
- Monitor only
- 4TTL (seconds)
- 3600
- 4Max payload bytes
- 262144
- 4Lookup timeout (ms)
- 500
04
Read the savings preview
The card shows a blast-radius preview built from your real traffic over the last 30 days. It estimates how many requests would repeat and what that would be worth.
| You see | Meaning |
|---|---|
| Repeated requests | How many requests in the window matched an earlier one |
| Estimated saving | What those repeats cost at your prices |
| Affected apps and models | Where the cache would act |
05
Choose how long answers live
Time to live is how long a stored answer may be reused. Pick it from how fast your answers go stale.
| Content | A good time to live |
|---|---|
| Static reference answers, embeddings of fixed text | 1 day to 30 days |
| Classification of repeated inputs | 1 hour to 1 day |
| Answers that depend on changing data | A few minutes, or do not cache |
The console starts at 3,600 seconds, one hour. The allowed range is one second to 30 days.
06
Set the size limit
Max payload bytes keeps very large responses out of the cache. The console starts at 262,144 bytes. A response above the limit is returned normally and not stored.
07
Switch to Enforce
When the preview and a few days of Observe look good, serve from cache.
- 1
Choose Mode: Enforce
Repeated requests are now answered from cache.
- 2
Start with one app
Do this on a policy bound to a single app first.
- 3
Watch the savings
The Dashboard's Realized Caching Savings card shows Cloptima exact cache savings.
To clear stored answers on demand, for example after you correct a wrong answer at the source, choose Invalidate cache on the policy's row in AI → Policies. Confirm, and every stored answer for that policy is dropped. The action is recorded in the audit log.
08
What makes two requests identical
A hit needs everything that shapes the answer to match.
| Part of the request | Must match? |
|---|---|
| The prompt and messages | Yes |
| The system prompt or instructions | Yes |
| Tools offered to the model | Yes |
| Settings such as temperature and top-p | Yes |
| The model and provider | Yes |
| Your customer, credential, and policy version | Always |
| App, team, and environment | By default. You can share across them |
| Caller labels, request ids, and user ids | No. They do not change the answer |
Answers are kept apart per customer and per credential. Another organization can never see them.
09
Recognize a hit
A response served from cache carries a header that says so.
HTTP/1.1 200 OK
content-type: application/json
x-cloptima-cache: hitUse it in tests and logs to confirm the cache is working. A request with no header went to the provider.
10
What is never cached
Some requests should always reach a provider. These are left out by default, and you choose how some of them behave.
- Streamed responses, unless you choose to cache them after they finish
- Responses larger than your size limit
- Requests that use tools the provider runs for you
- Requests that contain sensitive data, handled by your Bypass or Block choice
- Routes and models outside your cache lists
- Answers that were blocked or cut short by a guardrail
11
Keep it as code
Cache settings live in Terraform with the rest of the policy.
resource "cloptima_llm_gateway_policy" "classifier" {
name = "classifier"
mode = "enforce"
exact_cache_enabled = true
exact_cache_mode = "enforce"
exact_cache_ttl_seconds = 86400
exact_cache_max_payload_bytes = 262144
}12
If something goes wrong
Most surprises are about what counts as identical.
| What you see | Likely cause | Fix |
|---|---|---|
| No hits in Enforce | Requests differ in a setting such as temperature or a timestamp in the prompt | Make the prompt and settings stable, or use the semantic cache |
| Hits are low | The time to live is short, or traffic is mostly unique | Raise the time to live, and check the preview |
| A response is not cached | It was streamed, too large, or held sensitive data | Check the streaming and size settings |
| Old answers appear | The time to live is long for changing content | Shorten the time to live |