The number is only half the question#
100,000 reads a second of a small, cacheable page is easy — a CDN can serve it. 100,000 writes a second that each touch three tables is a hard project. Same number, very different work.
So ask five things first. Reads or writes? How big is each payload? What response time do you need — at p99, not on average? How many other services does one request call? Is the traffic steady, or does it spike?
Little’s Law sizes everything#
in flight = requests per second × response time
100,000 × 0.05 s = 5,000 requests in flight
100,000 × 0.20 s = 20,000 requests in flightOnly the response time changed — and now you must hold four times the connections, threads and memory. A slow dependency does not just slow you down. It multiplies what you have to keep open until something runs out.
What breaks, in order#
- The database connection pool. Response times jump off a cliff, not up a slope.
- One slow query or missing index — then it drags the shared pool down.
- Garbage-collection pauses. p99 goes spiky while the average looks fine.
- Slow services you call. You are fine; you are waiting on them.
- Network limits: ports, load balancer tables, TLS handshakes.
Keep the database pool small#
50 servers with 100 connections each is 5,000 connections. A Postgres server works best with a few dozen. More connections make it slower, not faster. Put a pooler such as PgBouncer in front, so many app connections share a few real ones.
Take work off the database#
- Cache at the CDN anything that is the same for everyone.
- Cache per user in Redis, with a version number in the key so an update makes the old value unreachable.
- Send reads that can be a second old to replicas.
- Do slow writes in the background: accept, save the request, return 202.
Stop the cache stampede#
If many keys were cached at the same moment, they also expire at the same moment. Every request misses at once and the database gets hit by all of them. Two fixes: add a little random time (jitter) to each expiry, and let only one request rebuild each key while the rest wait.
Plan for overload#
You will go over capacity one day. The question is what happens in that second. A system that accepts everything just gets slower until it falls over. A system that turns some requests away stays up.
- Bound every queue. An endless queue turns slowness into a crash.
- Reject early with 429 — before the request uses a database connection.
- Put a timeout on every outside call, and a circuit breaker in front of it.
- Make clients retry with backoff and jitter, not all at the same moment.
- Make writes idempotent. At this scale, retries of successful requests will happen.
Measure p99, not the average#
At 100,000 requests a second, “1% are slow” means 1,000 slow requests every second. And check that your load-testing tool corrects for coordinated omission — otherwise, when the system stalls, the tool quietly sends fewer requests and the stall never shows up in the results.