‹ All posts

How to test an LLM feature in CI

LLM answers change wording on every run, so normal tests break. Check structure, facts and meaning in layers, and pass on a percentage — that gives you a real gate.

LLM testingCI/CDEvaluation

The problem#

Run the same prompt twice and you get two different sentences. So assertEquals is useless, and many teams ship LLM features with no tests at all. The wording changes — but the facts you care about should not.

Test in three layers#

LayerChecksCost
StructureValid JSON, required fields, allowed valuesFree
FactsMust contain / must never containFree
MeaningTone, completeness, does it answer the questionOne model call

1. Structure#

Ask for JSON, then check it with code. The model saying “here is JSON” is not the same as valid JSON — a stray sentence before the opening brace is the most common break.

javascript
const answer = JSON.parse(await runFeature(testCase.input));
expect(validate(answer)).toBe(true);          // JSON Schema
expect(['approve', 'reject', 'escalate']).toContain(answer.decision);

2. Facts#

For each test case, write down what the answer must contain and what it must never contain. The “never” list matters more: it catches the moment the model starts citing a clause that exists — just not the right one.

3. Meaning#

For things code cannot check, use a second model as a judge. Keep it separate: it sees only the question, the answer and the evidence. And ask for a label — SUPPORTED or UNSUPPORTED — with one sentence of reason. A score out of ten changes every run; a label is steady.

Pass on a percentage#

LLMs sometimes fail a case they usually pass. If one red case fails the build, people start ignoring the tests. So pass the build when, say, 95% of cases pass — but structure failures and cases marked critical must always pass.

Keep the test set honest#

  • Write expected answers by hand or from the source — never by copying a run.
  • Delete cases nobody would bother to fix.
  • Turn every real bug from production into a new test case.
  • Record tokens and time in the same run, so a costly prompt change shows up too.