‹ All projects
DesignLLM testing

RAG retrieval tests

Tests the “find the right document” half of RAG on its own, because most bad answers start there.

The problem

When a RAG assistant answers badly, people blame the model. Usually the right document was never found.

How it works

  1. 01Question set
  2. 02Retrieve
  3. 03Recall@k
  4. 04Faithfulness
  5. 05Report

What was hard

  • Building test questions from real ones, not from what the search already returns.
  • An answer can match its source perfectly and still be wrong if the source is old.

The goal

Search and answer quality are scored separately, so effort goes to the half that is failing.

This is a design I worked out on paper — the problem, the approach and the trade-offs. It is not a shipped product.

Built with

Vector searchKeyword searchReranking