Two parts, tested separately#
RAG is two steps: search finds documents, then the model writes an answer from them. If you only test the final answer, you cannot tell which step failed.
| Search right? | Answer right? | What it means |
|---|---|---|
| Yes | Yes | Working. |
| Yes | No | A prompt or model problem. Now editing the prompt helps. |
| No | No | A search problem. Prompt changes will not help. |
| No | Yes | A lucky guess from the model’s memory. Dangerous. |
Measure recall@k#
For each test question, note which document has the answer. Then check whether it appears in the top k search results.
let found = 0;
for (const q of questions) {
const ids = (await search(q.text, k)).map((doc) => doc.id);
if (q.answerIds.some((id) => ids.includes(id))) found += 1;
}
const recallAtK = found / questions.length;Try k = 5 and k = 20. Good at 20 but poor at 5 means the right document is found but ranked low — add a reranker. Poor at both means it is not found at all — look at chunking and indexing.
Build honest test questions#
- Use real questions from users, tickets or search logs.
- Mark the right document by hand.
- Never use “whatever the search returns today” as the answer key — then the test only checks that search agrees with itself.
Chunking matters most#
- Too small, and a chunk has no context — a table row with no header.
- Too big, and the one useful sentence gets lost in the rest.
- Split on headings and sections, not a fixed number of characters.
- Keep the header row with every piece of a table.
Add keyword search#
Vector search is weak at exact codes like “CLM-2291” or “clause 4.2” — and those are what people search for. Run keyword search too, and merge the two result lists.
Watch for old documents#
An answer can match its source perfectly and still be wrong, because the source is out of date. Store a date with every chunk, prefer newer ones, and show the date next to the answer.