AI LAB / WHAT I INVESTIGATE
AI Lab
Experiments, benchmarks, failures, and reproducible conclusions.
EXP-014complete
RAG Retrieval on Local Models
RAG · Local AI · Benchmarks
- Recall@5
- 91.4%
- P95 latency
- 312ms
- Queries
- 500
EXP-015complete
Agent Tools Under Failure
Agents · Evaluation · Reliability
- Pass rate
- 86.4%
- Unsafe retries
- 2.1%
- Tests
- 120
EXP-016complete
Can LLMs Generate Schemas?
LLMs · Structured Output · Evaluation
- First-pass valid
- 89.8%
- After repair
- 97.5%
- Runs
- 80
EXP-017complete
Semantic Cache Thresholds
Caching · Embeddings · Performance
- Safe hit rate
- 41%
- False reuse
- 0.7%
- Pairs
- 900
EXP-018complete
Local Inference Routing
Local AI · Model Routing · Latency
- Local route
- 58%
- Quality delta
- -1.8%
- Cost saved
- 46%
EXP-019complete
Chunk Size vs Answer Quality
RAG · Chunking · Evaluation
- Best recall
- 94.1%
- Citation precision
- 90.7%
- Documents
- 180
EXP-020complete
Reranker Latency Budget
RAG · Reranking · Performance
- Best candidate count
- 20
- Added P95
- 184ms
- NDCG gain
- +11.2%
EXP-021inconclusive
Streaming JSON Repair
LLMs · Streaming · Structured Output
- Earlier useful data
- 420ms
- Repair success
- 72%
- Runs
- 200