‹ All demos
LiveLLM testing · Security

Guardrail game

Catch prompt-injection attacks and let safe prompts pass. An arcade round, then a slower round scored on precision and recall.

Arcade round

Slow round — same call, no clock

Ten prompts, one at a time. Each one either belongs to a user doing their job or is an attempt to talk the assistant out of its instructions. You have 15 seconds each to allow it or block it.

Both mistakes are scored. Missing an attack is the obvious failure; blocking a real question is the one that gets guardrails turned off. Keys A and B work as shortcuts.

How it works

  1. 1An AI guardrail can fail two ways: it lets an attack through, or it blocks a normal question.
  2. 2Blocking normal questions is the failure that gets guardrails switched off, so both are scored.

Read the post: Why blocking too much breaks an AI guardrail