← BACK TO PROJECT

PROJECT / Caching / Embeddings

Semantic Cache Service

A cache for expensive AI responses using embedding similarity, policy checks, and freshness controls.

CachingEmbeddingsPerformance
STATUS
shipped
YEAR
2026
ROLE
Backend Engineer
STACK
Python · Redis · PostgreSQL · FastAPI

SYSTEM / ARCHITECTURE

How it fits together

  1. 01Request normalizer
  2. 02Embedding index
  3. 03Policy engine
  4. 04Cache store
  5. 05Quality monitor
Cache hit rate
41%
Cost reduction
34%
P95 saved
1.8s
01

Constraint

Exact caches miss equivalent questions, but careless semantic reuse can return stale or unsafe answers.

  • Intent-sensitive matching
  • Tenant isolation
  • Freshness and policy controls
02

Approach

Candidates are retrieved by similarity, then checked against intent, scope, model version, and freshness before reuse.

  • Configurable thresholds
  • Negative cache decisions
  • Shadow-mode evaluation
03

Result

The service reduced cost and latency while keeping every reuse decision explainable in telemetry.