Overload is not a spike, it's a state
A service that can do 1,000 requests/second and receives 1,200 does not "run 20% slower". Queues fill, latency climbs past every client timeout, clients retry, and now it receives 2,000. Past saturation the useful throughput of an unprotected system drops towards zero — it spends all its time on requests that will be abandoned before they finish. Rate limiting (see Rate limiting) protects you from a single abusive client; backpressure and load shedding protect you from legitimate demand you cannot serve.
The mental model: every unbounded buffer is a promise you cannot keep. A queue with no limit, a thread pool that grows forever, a connection backlog of 10,000 — each one turns "we're overloaded" into "we're overloaded and every request now waits 40 seconds". Bounding the buffer forces the decision to the front, where refusing is cheap.
Backpressure: slow the producer
- Bounded queues everywhere: a work queue of 500, not infinity. When it is full the producer blocks or gets an immediate error, and the caller learns the truth now rather than after a timeout.
- Pull, don't push. Consumers that fetch the next batch when they are ready (Kafka-style) are self-regulating; a producer that pushes to a worker cannot know it is drowning. TCP flow control and reactive streams are the same idea at other layers.
- Concurrency limits over rate limits for your own tiers: cap in-flight requests to a dependency (say 64), and let the queue in front of that limit be short. Adaptive limiters (AIMD, gradient) find the number for you from latency.
Load shedding: refuse on purpose
When slowing the producer is not an option — a public API, a mobile app — you shed: reject a fraction of requests immediately with 429/503 and a Retry-After, so the rest complete within their deadline. Cheap rejection is the whole point; a shed request should cost microseconds at the edge, not a database round trip. Shed by priority: health checks and payments before analytics pings, logged-in users before crawlers, small requests before batch exports. Deadline propagation makes shedding smart: the caller sends the time it will give up (X-Request-Deadline), every hop subtracts its own cost, and any hop that sees an expired deadline drops the work instead of finishing an answer nobody is waiting for.
Diagram components: Client App, Load Balancer, Message Queue, Queue Worker, PostgreSQL.
The 503s are the design working: 900 useful responses instead of 1,200 timeouts.
Common mistakes
- Unbounded anything. A queue, pool or backlog with no ceiling converts overload into latency, and latency into retries — the amplification loop that turns a 20% overshoot into an outage.
- Shedding late. Rejecting after authentication, deserialisation and a database call means a shed request cost almost as much as a served one; refuse at the first hop that can.
- Random shedding. Dropping 20% of everything drops 20% of payments; rank the work and drop the cheapest-to-lose first.
- Retries without a deadline. A client that retries three times with a 30-second timeout each is a 90-second request in disguise; propagate one deadline and let every hop honour it.
Try it on a Grid
Open the Backpressure & Load Shedding challenge and focus on one thing: give every buffer a number, and label the edge where low-priority requests are refused with the status code and Retry-After you return. Queue-Based Load Leveling is the gentler cousin — same bounded queue, but the goal is smoothing rather than refusing — and Diagnose: Webhook Delivery Retry Storm shows what the amplification loop looks like from the inside.