The problem
When a RAG assistant answers badly, people blame the model. Usually the right document was never found.
How it works
- 01Question set
- 02Retrieve
- 03Recall@k
- 04Faithfulness
- 05Report
What was hard
- Building test questions from real ones, not from what the search already returns.
- An answer can match its source perfectly and still be wrong if the source is old.
The goal
Search and answer quality are scored separately, so effort goes to the half that is failing.
This is a design I worked out on paper — the problem, the approach and the trade-offs. It is not a shipped product.