Why Bigger Pools Wait Less
June 20, 2026
There is a result from queueing theory that feels wrong the first time you see it. Take a service behind a load balancer, keep every server exactly as busy as before, and just make the pool bigger. The latency gets better. Total throughput grows with the pool, but the less obvious benefit is lower latency at the same per-server utilization.
This is Marc Brooker's surprising economics of load-balanced systems. He holds each server's utilization fixed, grows the server count, and watches the queueing fall away. Most engineers already know the single-server version of the intuition from Erik Bernhardsson's "why you need to cut your systems some slack", where you run one machine below 80% to stay off the latency cliff. The multi-server story is the one that is less widely felt, so this post makes Brooker's result something you can drag a slider through, then takes it two steps further: when to merge pools instead of splitting them, and where real systems quietly hand the win back.
Everything below is computed live in your browser by a small open-source library I wrote for this post, queue-economics. Drag the sliders and the math reruns.
The model in one paragraph
Requests follow a Poisson arrival process, and service times are independent and exponentially distributed with a fixed mean. That mean is the service time, the unit for every wait number on this page: "0.5 service times" means half the average processing time. There are c identical servers, one shared first-come-first-served queue, and no timeouts or retries. The offered load stays below total capacity so the queue can reach a steady state. This is M/M/c. Under these assumptions, Erlang C gives the exact probability that an arrival has to wait. Real systems can depart from those assumptions, which is the subject of the last section.
The thing that makes a system feel slow is not utilization, it is waiting in line. And here is the trick: at a fixed utilization, the chance of having to wait drops sharply as the pool grows. With a handful of servers, one busy moment leaves an arrival nowhere to go. With many servers, there is almost always one free to grab the next request. More servers means more chances to dodge the queue.
Hold the utilization slider wherever you like and grow the pool. The wait probability falls off a cliff, and the p99 tail falls with it. If you do not fully trust a formula you met ninety seconds ago, good instinct, so flip on the Monte Carlo check: a seeded simulation drops its dots right on the analytic curve. Same sanity check Brooker ran on his version.
The "Size-biased wait" card answers a different question from the mean. Imagine sampling waiting requests in proportion to how long they wait, then measuring each one's total wait. Long waits get more weight, and requests that never queue get none. The result is E[W²] / E[W], or E[W] + Var(W)/E[W]. This is not the average wait of a randomly selected arrival.
For M/M/c, let C be the chance of waiting, λ the arrival rate, and μ each server's service rate. The ordinary mean wait is C / (cμ − λ). The size-biased mean is 2 / (cμ − λ): twice the mean among requests that actually wait, or 2/C times the mean across all arrivals. At the default 16 servers and 80% utilization, those displayed values are about 0.10 and 0.63 service times.
Brooker's Meet Alice. Alice is impatient. discusses the related inspection paradox. The remaining-wait quantity in that article is E[W²] / (2E[W]), half the size-biased total duration shown here. Total wait and remaining wait are different measurements.
The one rule worth remembering
If there is a single takeaway, it is this. To hit a target wait probability you staff roughly
servers ≈ load + β × √load
The bare load keeps the servers fed. The second term, the safety margin, is the slack you add so requests do not pile up. The surprise is that it grows like the square root of the load, not in proportion to it. This is the square-root staffing rule from the Halfin-Whitt regime.
A system handling 100x the traffic needs only about 10x the absolute slack, and proportionally far less. That is the whole economy of scale in one line: big systems can run hotter for the same quality of service, because their cushion is cheaper per unit of load. Small systems cannot, which is the uncomfortable corollary nobody mentions.
Split or merge?
This is where the theory turns into an architecture decision you actually make. Say you have some traffic and some servers. You can split them into several independent pools, or feed everything through one shared queue. Same total servers, same total traffic, same utilization either way. The only difference is whether the servers share a line.
Here is the intuition, because it is the part worth keeping. You only ever wait when every server is busy at the same instant. Split your servers into separate pools and each little pool hits that all busy state on its own, fairly often, while servers in a neighboring pool sit idle and cannot help. That idle time is stranded: real capacity that exists but is walled off from the request that needs it. Merge the pools and no server is ever idle while someone waits, because any free server anywhere takes the next request. It is the supermarket lesson: one line feeding ten checkouts beats ten separate lines, because you never get stuck behind one slow cart while a register opens up two aisles over. The bigger the shared pool, the rarer it is for every server to be busy at once, so the wait probability keeps falling as the pool grows. This is the square-root staffing rule from earlier seen from the other side: pool the load and one proportionally smaller cushion of spare servers covers everyone, instead of every small pool buying its own.
Within this shared-queue model, merging reduces waiting without adding servers. The wait probability drops and the tail shrinks, with no change to how hard each server works. The "cost of staying split" number is the concrete version: how many extra servers you would have to buy, across all your independent pools, just to match the merged pool's p99. That overhead is the price of fragmentation. It is also the argument for consolidating chatty microservices, sizing one big worker pool instead of many small ones, and being suspicious of per-tenant isolation that shards your capacity into pieces too small to benefit.
This is not a thought experiment. Around 2010 Heroku quietly changed its router from a single global request queue, which handed each request to a free dyno, to random routing, which drops each request onto one dyno's own queue. With single-threaded Rails dynos that turned one shared line into hundreds of independent ones, and the tail fell apart: a request could sit waiting while other dynos sat idle. It took James Somers' 2013 writeup for Heroku to admit the change, by which point Rap Genius was paying around twenty thousand dollars a month for the degraded throughput. Erik Bernhardsson compresses the whole lesson into one line: a 2x faster machine with a single queue always beats two 1x machines with their own queues. Sharing the queue is the entire game.
That Heroku split was accidental, which is why it was pure loss. Deliberate splits are a different story, and they are everywhere for good reason. The reason to split is almost never latency, it is failure. One shared queue is also one shared blast radius: a bad deploy, a poison request, or a single overloaded node can take the whole pool down at once. Availability zones, regions, and cells exist precisely so that one failure does not become every failure, and you cannot merge across that boundary without giving up the independence that protects you. Tenant isolation, compliance boundaries, and noisy-neighbor control carve up capacity for similar reasons. So queueing economics are one side of the ledger and redundancy is the other, and they pull in opposite directions. The move that respects both is to pool aggressively inside a failure domain and keep the domains separate: one healthy pool per zone beats ten small pools per zone, and it also beats a single global pool with nowhere to fail over to. Merge for latency and cost, split for blast radius, and put the boundary where a failure should stop, not where the org chart happens to fall.
Where the win leaks away
The clean curve assumes a tidy world: Poisson arrivals, independent exponential service times, infinite patience, and no retries. Real systems are not tidy, and the gap between the textbook and production is where most of the interesting failures live. Three distortions matter more than the rest.
Retries. A retry adds load at the worst possible moment, when the pool is already saturated. Here is the model behind the red line above, in words. Start with the offered load, the work actually arriving. Some fraction of requests have to wait, and the retry slider sets what share of those waiting requests give up and try again, piling extra work back on. So the effective load is the original load multiplied by (1 + retry rate × P(wait)). The trap is that P(wait) is itself driven by the load, so more load means more waiting, which means more retries, which means more load. The red curve is the fixed point of that loop: feed the new load back in, recompute, repeat until it stops moving. While the effective load stays under the server count, it settles on a slightly worse but stable number. Once retries push the effective load up to the number of servers, there is no stable point, the loop runs away, and the curve pins to 100%. That cliff is a real outage mode, the classic retry storm: a few timeouts trigger retries, the retries add load, the added load causes more timeouts, and a pool that was comfortable at 80% is suddenly at 100% and climbing. It is why retry budgets, exponential backoff, and circuit breakers exist, and why naive client retries turn a small wobble into a full outage.
Serialized shards. The flat line models each shard as an independent M/M/1 queue: one server, Poisson arrivals, and exponential service times. At the same per-shard utilization, adding shards adds total capacity but does not reduce the wait probability within a shard. Spare capacity in one shard cannot serve another shard's queue. A Durable Object is not automatically this model: its single-threaded JavaScript can interleave requests during asynchronous I/O. The approximation fits work that really is serialized, with the stated arrival and service-time assumptions.
Cold starts and fat tails. M/M/c assumes service times that are memoryless. Cold starts, GC pauses, and the occasional slow dependency give you a heavy tail instead, and the tail is exactly what p99 cares about. Pooling can still help, but Erlang C no longer gives an exact prediction. Measure the service-time distribution and use a suitable model or simulation for the tail.
So what
Let me step out of the server room for a second. The same math is why one big on-call rotation rides out a bad night better than three small ones, and why a team you keep pinned at 100% utilization always seems to have a mysteriously long queue of half-finished work. That is an analogy, not an exact M/M/c prediction: people differ in skills, tasks have dependencies, and coordination takes time. Shared capacity can help, but those constraints matter when applying the lesson to a team.
Back to machines. Three things to take away:
Bigger shared pools are not just more capacity, they are lower latency at the same cost, or the same latency at lower cost. The benefit shows up at modest sizes, not just at hyperscale, so you do not need to be Google to collect it. And the moment you shard, retry, or serialize, you start handing the win back, so spend it deliberately.
The math is old and settled. What is usually missing is a way to feel it and a number to put on the decision. That is what the sliders above are for, and the queue-economics package (on npm) is there if you want the same numbers in your own capacity planning.
If your system is sitting on the steep part of one of these curves and you want a second pair of eyes on it, book a call.