The problem
LLM features get tested for correctness, then shipped without anyone knowing how they behave under real traffic or what they cost.
How it works
- 01Scenario
- 02Ramp up
- 03Measure
- 04Cost model
- 05Threshold
What was hard
- Time to first token and total time are different numbers with different budgets.
- Telling a real slowdown apart from provider rate limits.
The goal
A known cost per user flow and a known concurrency limit, checked every release.
This is a design I worked out on paper — the problem, the approach and the trade-offs. It is not a shipped product.