Quantitative Methods

Queueing Theory and Why Utilization Above 85 Percent Breaks

Waiting time does not rise in proportion to load. It rises with one over the slack, which means the last few points of utilisation cost more than all the ones before them combined — and explains why the efficient department suddenly cannot hit a deadline.

  • Advanced
  • 12 min total
  • 14 chapters

What decision this helps you make: How much capacity to staff, and what utilisation target to set — with the waiting cost of the last point of utilisation priced rather than assumed away.

What this topic is

Queueing theory describes what happens when work arrives at random times and takes a variable amount of time to do. Its central result is that delay is driven by the product of three things: how full the system is, how variable the work and its arrivals are, and how long a single job takes. The utilisation term is non-linear — it is proportional to one over the remaining slack — so delay explodes near full capacity rather than degrading gracefully.

Why it matters

Utilisation is the most commonly reported efficiency metric in operations, and it is a metric whose improvement buys you a hidden, accelerating cost. Managers push a team from 80% to 92% loaded, report a productivity gain, and are then surprised when lead times triple and expediting becomes a full-time job. The arithmetic here tells you in advance what the last point of utilisation costs in delay, so the trade can be made deliberately rather than discovered.

Who should learn it

Anyone staffing a service operation, sizing a support or claims team, planning machine loading, running a professional services firm, or defending a capacity request against a utilisation target.

What you will understand

  • Why waiting time scales with one over the slack, and what that does between 80% and 95% utilisation
  • How to size a multi-server team so that a specific waiting-time target is met, and what one more server buys
  • Why variability and utilisation are substitutes — you can buy the same delay reduction from either
  • Where high utilisation is genuinely correct, and why there is nothing magic about the number 85

Prerequisites

Common misconception

"We are only at 90% capacity, so we have 10% headroom." You do not have 10% headroom; you have a system whose average queue is roughly nine times its service time and whose delay doubles again by 95%. The linear intuition — 90% full means 90% of the delay of 100% full — is exactly wrong, because the relationship is a reciprocal, not a proportion. Between 80% and 90% utilisation, average waiting time in a simple single-server queue does not rise by an eighth. It rises by a factor of well over two.