PROJECT / Caching / Embeddings
Semantic Cache Service
A cache for expensive AI responses using embedding similarity, policy checks, and freshness controls.
CachingEmbeddingsPerformance
- STATUS
- shipped
- YEAR
- 2026
- ROLE
- Backend Engineer
- STACK
- Python · Redis · PostgreSQL · FastAPI
SYSTEM / ARCHITECTURE
How it fits together
- 01Request normalizer
- 02Embedding index
- 03Policy engine
- 04Cache store
- 05Quality monitor
- Cache hit rate
- 41%
- Cost reduction
- 34%
- P95 saved
- 1.8s
Constraint
Exact caches miss equivalent questions, but careless semantic reuse can return stale or unsafe answers.
- Intent-sensitive matching
- Tenant isolation
- Freshness and policy controls
Approach
Candidates are retrieved by similarity, then checked against intent, scope, model version, and freshness before reuse.
- Configurable thresholds
- Negative cache decisions
- Shadow-mode evaluation
Result
The service reduced cost and latency while keeping every reuse decision explainable in telemetry.