There is a number that most organizations chase without ever questioning it: resource utilization, usually reported as a utilization rate. It is the percentage of available time that people spend working on assigned tasks. High utilization feels like efficiency. A team running at ninety-five percent looks lean, disciplined and well managed. A team running at seventy percent looks like slack that should be filled.
This instinct is one of the most expensive in portfolio management. On the resource that actually limits your delivery, pushing utilization towards a hundred percent does not make the portfolio faster. It makes it dramatically slower. And the reason is not effort or discipline. It is arithmetic, from the same branch of mathematics that explains why motorways seize up before they are physically full.
A utilization rate is not a bad measurement. It tells you, accurately, how busy people are. Where it goes wrong is as a target, because it treats every hour in the organization as equally important. Delivery speed is not set by every hour. It is set by the constraint: the one team, skill or person whose capacity limits how fast the whole portfolio can deliver. So Flow Economics asks two different questions. How loaded is the constraint, and how long does work take to move from start to delivery? That second number is flow time, and by the end of this piece it should look like the one worth reporting.
Why a road jams before it is full
A motorway can carry a certain number of cars per hour. You might think it runs smoothly right up until it is completely full, then stops. It does not. Traffic starts to break down while there is still visible road between vehicles, at something like eighty-five to ninety percent of theoretical capacity. Past that point, a tiny disturbance, one driver braking, causes a wave that everyone behind has to absorb, because there is no spare room for the system to recover. The closer you run to full, the longer it takes to clear each disturbance, and the delays grow out of all proportion to the extra traffic.
A constrained team is a road, and tasks are the cars. Below a certain load, work flows and queues clear quickly. Above it, every new arrival waits behind a queue that no longer has time to empty, and waiting time climbs steeply while utilization barely moves. The system is not broken. It is doing exactly what queueing systems do.
Why does a high utilization rate slow everything down?
Queueing theory gives this a precise shape. For a resource fed by variable demand, the average time work spends waiting is proportional to a simple expression: utilization divided by one minus utilization.
Put numbers to it. At fifty percent utilization, that figure is one. At eighty percent it is four. At ninety percent it is nine. At ninety-five percent it is nineteen. The move from fifty to ninety percent utilization, which on a capacity report looks like sensible tightening, multiplies the average wait by nine. The move from ninety to ninety-five percent, which looks like a rounding error, roughly doubles it again.
This is the part managers rarely see, because the two numbers they can measure move in opposite directions and at completely different speeds. Utilization crawls from ninety to ninety-five. Delivery time lurches. On the dashboard the team looks marginally more efficient. In reality every project touching that team just got materially later.
And this curve assumes an average, well-behaved workload. Real work is not well behaved. Estimates are wrong, scope shifts, urgent requests arrive, people take leave. That variability is a multiplier on the whole curve: the more uneven the work and the arrivals, the higher the waiting time at every level of utilization. High load and high variability together are what turn a busy team into a stalled one.
Protective capacity is not waste
The uncomfortable conclusion is that a constrained resource needs a margin of unused time to function, and that margin is not slack to be eliminated. It is what lets the queue clear between disturbances. Take it away and the queue never recovers; it just grows, and everything downstream waits.
This spare margin has a name in Theory of Constraints: protective capacity. It is the reserve on the constraint that absorbs variability and keeps work flowing. Running the constraint at a hundred percent does not capture that reserve as free output. It converts it into an ever-lengthening queue of half-finished work, which is the most expensive form of inventory an organization can hold, because it has consumed capacity and returned nothing.
So the reserve is doing real work even when it looks idle. It is buying speed and predictability for everything that passes through the constraint. Costing it as waste is the same accounting error that treats a started project as progress and a queue as productivity.
Where this actually bites
None of this means you should run the whole organization at seventy percent. Most resources are not the constraint, and idle time on a non-constraint is genuinely cheap. The utilization trap is specific: it bites hardest on the one team, skill or individual that governs the pace of delivery.
That is the resource everyone is waiting on. Load it to the limit and its queue explodes, and because every project routes through it, the delay is shared across the entire portfolio. Load a non-constraint resource to the limit and, at worst, it produces work that then sits waiting for the constraint anyway. The whole system can only move at the speed of its slowest necessary step, so the only utilization figure that decides portfolio speed is the utilization of the constraint. Chasing high utilization everywhere else optimises numbers that do not govern the outcome, a classic case of local efficiency bought at the cost of global throughput.
This is why the reflex to keep everyone busy is not just harmless bookkeeping. It actively loads the constraint, because the easiest way to make a team look fully utilized is to start more work and push it into the system, precisely the behavior that drives the constraint up its waiting curve. The pursuit of local efficiency manufactures the global delay.
What should you measure instead of utilization rate?
The fix is to stop treating utilization as a target and start treating it as a dial with a correct setting that is not a hundred percent.
First, identify the constraint honestly, and manage its load deliberately. The goal is to keep it highly productive on the right work, not maximally busy on all work. That means protecting a margin of reserve capacity on it, and defending that margin when someone points at it and calls it waste. Reserve on the constraint is not the place to find efficiency. It is the place efficiency comes from.
Second, control what you feed it. A constraint climbs its waiting curve because too much work is pushed in, so the lever is to cap what is in flight and release new work only as the constraint frees up. This is the same discipline as holding a portfolio work-in-progress limit: less work in the system means lower effective utilization on the constraint, shorter queues, and faster finishes on the same capacity. And when a slot on the constraint does open, value per constrained resource hour is what decides which work earns it.
Third, change the number you report. Utilization answers "how busy are people," which is not the question that matters. The question that matters is "how fast does work move through the system." Measure flow time, the elapsed time from when work starts to when it delivers, and watch what happens to it as constraint load changes. It will tell you the truth that utilization hides: that past a certain point, loading the constraint harder makes everything slower.
The shift
Full utilization is intuitive, easy to measure, and wrong as a goal. It rewards the exact behavior, loading the constraint to the limit, that lengthens every queue in the portfolio, and it flatters a dashboard that reads as efficient while delivery quietly slows. This is another face of the coordination ceiling: the point past which doing more at once makes the organization deliver less.
A portfolio is not a machine that runs best when every part is redlined. It is a flow system, and flow systems run fastest with a margin to spare on the parts that constrain them. The counter-intuitive move is to deliberately leave room on your most valuable resource, protect that room, and judge the portfolio by how fast value reaches the finish line rather than by how fully your people are booked. Run the constraint a little below the limit, and the whole system speeds up.
—
Flow Economics is the discipline of understanding how value actually moves through an organization when its resources are constrained, and of making decisions that maximize what the whole system delivers rather than what each part looks like on its own. If your portfolio is measured by how busy its people are rather than how fast its work finishes, the Flow Economics framework is a good place to see the full picture.

Leave a Reply