‹ All projects
DesignLLM flows

LLM gateway

One endpoint in front of several models: routing, fallback, caching and a spend limit per team.

The problem

Every prototype called a model provider directly, with its own key and retry logic. Nobody knew what a feature cost.

How it works

  1. 01Request
  2. 02Route
  3. 03Cache
  4. 04Call + fallback
  5. 05Meter cost

What was hard

  • Switching providers mid-stream without repeating text the user already saw.
  • Cache keys include the prompt version, so a prompt change is not hidden by old answers.
  • Hitting the budget moves to a cheaper model instead of failing.

The goal

Cost and speed per feature in one place, and an outage at one provider slows things down instead of breaking them.

This is a design I worked out on paper — the problem, the approach and the trade-offs. It is not a shipped product.

Built with

StreamingCacheCircuit breaker