AI LAB / RAG / Reranking
Reranker Latency Budget
Measuring how many candidates a reranker can score before response latency stops being worthwhile.
RAGRerankingPerformance
- EXPERIMENT
- EXP-020
- STATUS
- complete
- DATE
- Jul 22, 2026
- TAGS
- RAG · Reranking · Performance
EXPERIMENT / HYPOTHESIS
Reranking 20 candidates gives most of the quality gain within a 250ms interactive latency budget.
- 01Retrieve 10, 20, 40, and 80 candidates
- 02Measure NDCG and end-to-end latency
- 03Repeat across short and long documents
- Best candidate count
- 20
- Added P95
- 184ms
- NDCG gain
- +11.2%
CONCLUSIONTwenty candidates captured most quality improvement. Larger sets produced diminishing gains and unstable tail latency.
Budget
The experiment assigned 250ms of an 800ms retrieval budget to reranking, including serialization and network overhead.
- Warm and cold runs
- Batch sizes
- Tail latency
Decision
The production default is 20 candidates, with lower counts for simple queries and asynchronous reranking for research mode.
- Adaptive count
- Budget-aware fallback
- Observe P95