The problem
LLM output changes wording every run, so normal “equals” tests fail — and teams end up with no tests at all.
How it works
- 01Golden set
- 02Run
- 03Check
- 04Compare
- 05Gate
What was hard
- Check facts and structure, not exact wording.
- Pass on a percentage of cases, with some cases that must always pass.
- Expected answers written by a person, never copied from a run.
The goal
A prompt change that hurts accuracy fails the build, with the exact cases it broke.
This is a design I worked out on paper — the problem, the approach and the trade-offs. It is not a shipped product.