Semantic Cache

Reuse What Your LLM Traffic Already Answered by Meaning

Cloptima matches requests semantically to serve near-duplicate requests from a local vector cache that activates only where evals, freshness, and similarity thresholds say it's safe.

app.cloptima.ai/llm/policies
Semantic response cache
Namespace: support-ai · similarity threshold: 0.92
Illustrative
Can I get a refund on an annual plan? — 0.95 similarity
model=gpt-4o · freshness: 4h · eval pass
Cache hit (enforce)
How do annual subscription refunds work? — 0.93 similarity
model=gpt-4o · freshness: 2h · eval pass
Cache hit (enforce)
What is your refund policy for enterprise? — 0.84 similarity
similarity below threshold · forwarded to provider
Cache miss
Semantic cache hit rate (observe)
19%
Eval gate status
Passed (0.94)

Exact match misses most of the repeat traffic

Exact-match caching only catches requests that are byte-for-byte identical, but a large share of repeat traffic is semantically identical: the same question asked in different words, or the same classification with different phrasing. Without semantic caching, that traffic re-runs the full model every time, driving up latency and provider token spend.

  • Byte-identical caching misses the larger population of near-duplicate requests
  • Identical user intents worded slightly differently incur full model latency and cost
  • Wrong-answer risk from serving a semantically 'close enough' response with no safety gate
  • Risk of unverified or low-fidelity vector representations serving cached responses

A strict safety gate on top of vector similarity

Cloptima matches requests semantically using local vector embeddings against previous responses. Semantic cache eligibility is a strict gate: a request only qualifies for a cached hit after passing namespace isolation, route and model-family matching, tool-schema matching, source-freshness checks, and content-class checks — and, for enforce mode, an eval-score threshold and optional human approval. Enforce mode requires verified production embeddings and passing eval thresholds before live serving begins.

  • Near-instant local vector matching across semantic candidate entries
  • Semantic cache observe mode: log would-be matches and similarity scores with zero behavior change
  • Semantic cache enforce mode: serve only after namespace, route, model-family, tool-schema, freshness, and content-class checks pass
  • Eval-score gate and optional human-approval gate before enforce-mode serving
  • Strict verification gate — high-fidelity embeddings and eval passes are required before enforce mode can activate

Observe first, enforce second

Start with semantic cache observe mode once embeddings are configured. Review similarity scores, match quality, and projected cost savings with zero traffic disruption. Reserve enforce mode for narrow, eval-scored request scopes once confidence is proven.

Outside the synchronous provider path

Semantic cache candidate lookups run entirely outside the synchronous provider request path — a candidate lookup either returns fast or the request proceeds to the provider with no added latency floor. The candidate endpoint never returns response bodies directly, only pointers (cache key hash and policy id), so the gateway engine replays a hit through the same safety path as the exact-cache layer rather than trusting raw vector proximity.

Launch path

Configure a BYOK embedding credential for your semantic cache runtime. Turn on semantic cache observe mode to review match quality and simulated hit rate. Require a passing eval score before moving any policy scope to enforce mode.

FAQ

Operationalize LLM FinOps Across Your Apps

Start with telemetry, gateway governance, or provider bill matching workflows. Keep model spend connected to engineering ownership and finance reporting.

No credit card required
5-minute setup
Free trial