AI LAB / Caching / Embeddings
Semantic Cache Thresholds
Finding a similarity threshold that saves cost without returning incorrect cached answers.
CachingEmbeddingsPerformance
- EXPERIMENT
- EXP-017
- STATUS
- complete
- DATE
- Aug 12, 2026
- TAGS
- Caching · Embeddings · Performance
EXPERIMENT / HYPOTHESIS
A threshold calibrated by intent category can outperform one global similarity cutoff.
- 01Label 900 query pairs as reusable or unsafe
- 02Sweep global thresholds
- 03Compare against per-intent calibration
- Safe hit rate
- 41%
- False reuse
- 0.7%
- Pairs
- 900
CONCLUSIONPer-intent thresholds produced more safe hits, especially for factual lookup, while transactional queries required near-exact matches.
Dataset
Pairs included paraphrases, changed entities, time-sensitive questions, and superficially similar requests with different intents.
- Safe reuse
- Unsafe reuse
- Ambiguous
Result
Similarity alone was insufficient. Intent and freshness policy were necessary features in the reuse decision.
- Calibrate per intent
- Exclude volatile data
- Audit sampled hits