← BACK TO AI LAB

AI LAB / Caching / Embeddings

Semantic Cache Thresholds

Finding a similarity threshold that saves cost without returning incorrect cached answers.

CachingEmbeddingsPerformance
EXPERIMENT
EXP-017
STATUS
complete
DATE
Aug 12, 2026
TAGS
Caching · Embeddings · Performance

EXPERIMENT / HYPOTHESIS

A threshold calibrated by intent category can outperform one global similarity cutoff.

  1. 01Label 900 query pairs as reusable or unsafe
  2. 02Sweep global thresholds
  3. 03Compare against per-intent calibration
Safe hit rate
41%
False reuse
0.7%
Pairs
900

CONCLUSIONPer-intent thresholds produced more safe hits, especially for factual lookup, while transactional queries required near-exact matches.

01

Dataset

Pairs included paraphrases, changed entities, time-sensitive questions, and superficially similar requests with different intents.

  • Safe reuse
  • Unsafe reuse
  • Ambiguous
02

Result

Similarity alone was insufficient. Intent and freshness policy were necessary features in the reuse decision.

  • Calibrate per intent
  • Exclude volatile data
  • Audit sampled hits