Backpressure & load shedding

What to do when work arrives faster than you can finish it: bound the queues, refuse early, keep the important work, and make deadlines travel with the request.

Overload is not a spike, it's a state

A service that can do 1,000 requests/second and receives 1,200 does not "run 20% slower". Queues fill, latency climbs past every client timeout, clients retry, and now it receives 2,000. Past saturation the useful throughput of an unprotected system drops towards zero — it spends all its time on requests that will be abandoned before they finish. Rate limiting (see Rate limiting) protects you from a single abusive client; backpressure and load shedding protect you from legitimate demand you cannot serve.

The mental model: every unbounded buffer is a promise you cannot keep. A queue with no limit, a thread pool that grows forever, a connection backlog of 10,000 — each one turns "we're overloaded" into "we're overloaded and every request now waits 40 seconds". Bounding the buffer forces the decision to the front, where refusing is cheap.

Backpressure: slow the producer

  • Bounded queues everywhere: a work queue of 500, not infinity. When it is full the producer blocks or gets an immediate error, and the caller learns the truth now rather than after a timeout.
  • Pull, don't push. Consumers that fetch the next batch when they are ready (Kafka-style) are self-regulating; a producer that pushes to a worker cannot know it is drowning. TCP flow control and reactive streams are the same idea at other layers.
  • Concurrency limits over rate limits for your own tiers: cap in-flight requests to a dependency (say 64), and let the queue in front of that limit be short. Adaptive limiters (AIMD, gradient) find the number for you from latency.

Load shedding: refuse on purpose

When slowing the producer is not an option — a public API, a mobile app — you shed: reject a fraction of requests immediately with 429/503 and a Retry-After, so the rest complete within their deadline. Cheap rejection is the whole point; a shed request should cost microseconds at the edge, not a database round trip. Shed by priority: health checks and payments before analytics pings, logged-in users before crawlers, small requests before batch exports. Deadline propagation makes shedding smart: the caller sends the time it will give up (X-Request-Deadline), every hop subtracts its own cost, and any hop that sees an expired deadline drops the work instead of finishing an answer nobody is waiting for.

Bounded queue in front, priority shedding at the edge, deadline travelling with the request

Diagram components: Client App, Load Balancer, Message Queue, Queue Worker, PostgreSQL.

The 503s are the design working: 900 useful responses instead of 1,200 timeouts.

Note Circuit breaker vs shedding. A breaker protects the caller from a sick dependency (stop calling it, fail fast, probe later). Shedding protects the callee from too many healthy callers. You want both: the breaker in every client, the shedder in every server — and neither replaces the other.

Common mistakes

  • Unbounded anything. A queue, pool or backlog with no ceiling converts overload into latency, and latency into retries — the amplification loop that turns a 20% overshoot into an outage.
  • Shedding late. Rejecting after authentication, deserialisation and a database call means a shed request cost almost as much as a served one; refuse at the first hop that can.
  • Random shedding. Dropping 20% of everything drops 20% of payments; rank the work and drop the cheapest-to-lose first.
  • Retries without a deadline. A client that retries three times with a 30-second timeout each is a 90-second request in disguise; propagate one deadline and let every hop honour it.

Try it on a Grid

Open the Backpressure & Load Shedding challenge and focus on one thing: give every buffer a number, and label the edge where low-priority requests are refused with the status code and Retry-After you return. Queue-Based Load Leveling is the gentler cousin — same bounded queue, but the goal is smoothing rather than refusing — and Diagnose: Webhook Delivery Retry Storm shows what the amplification loop looks like from the inside.

All Learn articles