The Variability Tax: Why Smoothing Your Work Buys More Speed Than Adding People

Two organizations run the same specialist gate. Both have the same number of people in it, both push the same volume of work through it in a year, and both, when you ask the capacity planning team, report the same utilization figure: about 85 per cent. On any dashboard either board has ever seen, these two organizations are identical.

In the first, work waits about a week at the gate. In the second, it waits about eight weeks.

Nothing in the capacity model explains that gap, because the thing causing it is not in the capacity model. It is not headcount, not demand, not skill, not effort, and not management grip. It is unevenness: how lumpy the arrivals are, and how variable the work items are once they arrive. Unevenness has a price, that price is paid in delay, and delay is money at a rate you can calculate. I want to give it a name, because an unnamed cost gets absorbed rather than managed: the variability tax.

The reason this matters more than it sounds is the direction of the lever. Most portfolios, confronted with an eight-week queue, reach for capacity: hire, borrow, outsource, work weekends. Smoothing is almost always cheaper, almost always faster, and in the arithmetic below it buys about sixty per cent of the delay back for no additional headcount at all.

The two dials, not one

The Utilization Trap taught the first dial. Waiting time is not proportional to how busy you are, it is proportional to utilization over one minus utilization, so the curve bends upward and then goes vertical as you approach a full constraint. That is the dial almost everyone in portfolio management now knows about.

There is a second dial sitting right beside it in the same equation, and it gets a fraction of the attention. The standard approximation for waiting time in a queue has three terms multiplied together:

Average wait = (a variability term) x (a utilization term) x (the average time the work itself takes).

The utilization term is the familiar one. The variability term is the average of two squared coefficients of variation: one for how unevenly work arrives, one for how unevenly long each item takes. A coefficient of variation is just the standard deviation divided by the mean, so it is a pure measure of spread with the size scaled out. Arrivals that trickle in at a steady rhythm score around 0.2 or 0.3. Arrivals that come in clumps, twelve one week and none the next, score above 1.

Here is the part that changes decisions. Those two terms multiply. They do not add, they do not average, and neither one caps the other. Which means a portfolio can be only moderately loaded and still slow, if it is lumpy enough. And it means that at high utilization, where most constraints actually run, a modest increase in unevenness produces an enormous increase in waiting, because it is being amplified by a utilization term that is already large.

Utilization and variability are not two competing explanations for your queue. They are two factors of one product, and you are almost certainly managing only one of them.

Where the unevenness actually comes from

Ask most delivery leaders why their work is uneven and you will get a version of "because the world is uneven". Some of it is. Most of it is manufactured inside the organization, by decisions taken for unrelated reasons, and once you can see the five main sources you can see how much of the tax is self-imposed.

Batched arrivals from your own planning cadence. If the portfolio commits work once a year, work does not arrive at the constraint smoothly across that year. It arrives in a wave in the weeks after approval, because every project that was funded starts at once and every project that starts needs the same specialists first. Then there is a second wave before year end, when everyone rushes to spend or to close. This is the frozen portfolio seen from the constraint's point of view: the annual cycle is not just a bad information decision, it is a variability generator, and it is one your organization designed and can redesign.

Unplanned operational demand. Incidents, escalations, regulatory requests, warranty work on last quarter's delivery. Phantom capacity covers what this does to the average: it removes hours the plan thought it had. Its effect on variability is separate and often larger. Unplanned work is not merely unbudgeted, it is unbudgeted and bursty, and it lands preferentially on exactly the scarce people whose queue is already the most sensitive to disruption.

Expedites. Every time a piece of work is pushed to the front, two things happen. The obvious one is that everything behind it waits longer, which is the expedite tax. The less obvious one is that the expedite makes the whole queue less predictable from then on, for every item in it, because an item's wait now depends on how many interruptions happen to arrive during it. Ten expedites a quarter do not just cost ten displacements. They raise the variance for all of the several hundred items that were not expedited.

Work items of wildly different sizes. If your gate handles a two-day review and a nine-week review through the same queue, the service-time coefficient of variation is high before anyone has done anything wrong, and the nine-week item blocks everything behind it in a way no amount of prioritization undoes. This is the queueing consequence of batch size, and it is why splitting large items helps even when the total work is unchanged.

Rework loops. An item that comes back is an arrival you did not forecast, at a time you cannot predict, of a size you do not know until you look. The rework spiral costs you the redone effort, which everybody counts. It also injects variance into the arrival stream, which nobody counts, and that second cost falls on every other project sharing the gate.

Notice what these five have in common. Not one of them is a fact about your industry, your customers, or the inherent unpredictability of knowledge work. Every one is a consequence of a policy: how often you plan, where unplanned work is allowed to land, whether expediting is free, how big you let items get, and how much you invest in getting things right first time. The variability tax is mostly a bill for policies, and bills for policies can be renegotiated.

A worked example

Numbers make this concrete. These are illustrative rather than precise, and the queue is modeled as a single stream to keep the arithmetic legible; a gate staffed by several people softens the effect without removing it. The pattern is what matters.

A portfolio runs everything through a shared validation and assurance gate. An item takes four working days of gate time on average. The gate runs at 85 per cent utilization, which the capacity team is rather proud of, since it means the specialists are neither idle nor visibly drowning. At that load the utilization term is 0.85 divided by 0.15, which is 5.67.

Now take the two organizations from the opening.

The smooth one. Work is released to the gate on a weekly cadence, items are sized to a similar range, and unplanned demand is handled elsewhere. Coefficients of variation of about 0.5 on both arrivals and service, so the variability term is 0.25. Average wait: 5.67 x 0.25 x 4 = 5.7 working days. Total time at the gate, including the work itself, is about two working weeks.

The lumpy one. Same people, same throughput, same 85 per cent. But work arrives in post-approval waves, the gate absorbs incident escalations directly, expedites run at a few a month, and items range from two days to two months. Coefficients of variation of about 1.4 on arrivals and 1.3 on service, so the variability term is 1.825. Average wait: 5.67 x 1.825 x 4 = 41.4 working days.

Same capacity. Same workload. Same utilization. Seven times the wait, and a queue that is now longer than eight working weeks.

Two curves of average waiting time plotted against utilisation. The lower curve, for smooth arrivals and even work sizes, stays under six days all the way to 85 per cent utilisation. The upper curve, for lumpy arrivals and uneven work sizes, reaches 41 days at that same 85 per cent. A marker shows that halving the unevenness brings the wait down to 16 days without changing capacity, and a horizontal line shows that matching the smooth queue's wait by adding capacity alone would mean running the gate at 44 per cent utilisation.
Two organizations, one point on the horizontal axis, two entirely different delivery experiences. The vertical gap between the curves is the variability tax.

Put money on it. The difference is 35.7 working days of queue, about seven working weeks, every time an item passes the gate. A typical project in this portfolio passes twice, once at design assurance and once before release, so the lumpy organization adds roughly 71 working days, call it three and a quarter calendar months, to every project's elapsed life.

At this gate's throughput, about 26 projects complete a year. If a delivered project returns $1.2m a year once live, its drag cost is about $100,000 a month, so three and a quarter months of extra queue is $325,000 per project. Across the portfolio that is roughly $8.5m a year, paid entirely for unevenness, by an organization whose capacity plan says it has exactly the same resources as the one that is not paying it.

Now the two ways out, and the reason this article exists.

Buy your way out with capacity. To get the lumpy organization's wait down to the smooth organization's 5.7 days by adding people alone, leaving the unevenness exactly as it is, you would have to drive utilization down to about 44 per cent. That is not a trim. At constant demand it is very nearly doubling the gate. You would be recruiting an entire second team, and paying the whole of the hiring dip to get it, so that the queue could go on being lumpy in comfort.

Halve the unevenness instead. Suppose you do the unglamorous work: release to the gate weekly rather than in waves, put a triage layer in front of the unplanned demand, cap expedites, split the largest items. You do not achieve the smooth organization's discipline, you get halfway. Coefficients of variation fall from 1.4 and 1.3 to about 0.8 and 0.9, so the variability term drops from 1.825 to 0.725. Average wait: 5.67 x 0.725 x 4 = 16.4 working days.

That is 25 working days off every pass, 50 off every project, about two and a third calendar months, worth roughly $230,000 a project and close to $6m a year. No new headcount. No ramp. No requisition. Mostly a change to when things are allowed to start and where interruptions are allowed to land.

Sixty per cent of the delay, removed by changing the rhythm rather than the resources. That ratio is the practical content of this whole idea.

Why single-project management is blind to this

If the effect is this large, why does it not show up anywhere? Because every instrument the portfolio uses reports the mean, and variability is by definition the thing a mean deletes.

A capacity model works in averages by construction. Average demand, average capacity, average duration. Feed it two portfolios with identical averages and different spreads and it returns identical answers, confidently, because the spread was discarded at the point of data entry. The model is not making an error. It is answering a question about averages and being read as though it had answered a question about experience.

A single project sees only its own draw. It has one item at the gate, it waits nine weeks, and it records a nine-week wait as a fact about the gate. It cannot see that three of those weeks were the tail of the annual approval wave, two were someone else's expedite, and two more were a single enormous item ahead of it that should have been split. It has no visibility of the queue's composition, so it does the only thing available to it and adds contingency next time. Every project does the same, and the portfolio has now converted a variability problem into padding that hides the signal without touching the cause.

The cost also fails the attribution test in both directions. The part of the organization that creates a variability spike, by releasing thirty approved projects into the system in April, or by escalating an incident straight onto a specialist, never sees the bill, because the bill arrives weeks later and is spread thinly across every unrelated project sharing that gate. And the projects that pay it cannot trace their delay to a cause outside their own boundary. A cost with no visible payer and no traceable causer is a cost that no governance process will ever be asked to approve.

Underneath all of it sits the Project Illusion in its scheduling form: the belief that a project's duration is a property of the project. It is not. It is a property of the project and the queueing behavior of everything it shares capacity with, and a large part of a typical project's calendar is determined by decisions taken in projects its manager has never heard of.

What to do instead

1. Measure the spread, and put it next to the average. You almost certainly report average time at each gate. Report three numbers instead: the median wait, the 85th percentile wait, and the ratio between them. That ratio is your variability tax made visible, and it is the only one of the three that tells you whether the problem is load or unevenness. A gate with a five-day median and a forty-day 85th percentile has a variability problem that no amount of capacity will fix cheaply. Track the ratio monthly. It moves when you change policy, which makes it the rare portfolio metric that is genuinely a feedback loop.

2. Fix the arrival rhythm before you touch capacity. Arrival variability is usually the largest single term and the cheapest to change, because it is entirely under your control. Move from annual or event-driven release of work to a fixed cadence: a standing fortnightly release of approved work into the constraint's queue, in a quantity matched to what it can absorb, rather than everything the moment funding clears. This costs nothing and requires no new tooling. It is the same underlying move as limiting work in progress, applied to the tempo of starts rather than to their number, and the two reinforce each other.

3. Give unplanned work its own lane, and never let it hit the constraint's queue directly. The point is not to refuse operational demand, which is real and legitimate. The point is to stop it propagating its variance into planned work. Nominate a first-response capability that is deliberately not the constraint, resource it with the explicit expectation that it will sometimes be idle, and let it absorb, triage and resolve most of what arrives. Only genuine constraint work gets escalated through, and it gets scheduled rather than injected. A shock absorber does not remove the bump. It stops the bump reaching the passengers.

4. Price expedites, cap them, and publish the count. Make the expedite lane explicit and finite: a fixed number of slots per period, with every use of one requiring the displaced projects to be named. The cap matters more than the price. An organization that allows five expedites a quarter has a queue people can plan around. One that allows unlimited expedites has no queue at all, only a standing negotiation, and its variability term will never come down whatever else it fixes.

5. Compress the range of item sizes at the gate. Set a maximum size for anything entering the constraint's queue, and split what exceeds it, even when splitting adds a little total work. A gate whose items run from two days to two months will always be unpredictable. One whose items run from three to eight days is close to smooth by construction, and that smoothing is worth more than the handful of extra handovers costs. Where a split is genuinely impossible, schedule the outsized item explicitly rather than letting it queue like everything else, so that it disrupts on a known date instead of a random one.

6. Size protective capacity to your variability, not to a habit. Every organization has an unwritten target utilization for its specialists, usually somewhere between 85 and 95 per cent, and it is almost never derived from anything. Derive it. The higher your variability term, the more slack the same delivery performance requires, because the two multiply. That is an argument you can actually take to a finance director: not "the team needs breathing room", which sounds like a preference, but "at our current unevenness, 85 per cent utilization buys us a forty-day queue and 75 per cent buys us a twenty-two-day one, and here is what those eighteen days are worth at our value per constrained resource hour".

7. Treat every variability reduction as a capacity purchase, and compare it as one. When you smooth arrivals, or move the unplanned somewhere else, you have not saved time in some soft, hard-to-quantify sense. You have bought delivery speed, of the same kind and in the same units that a hire would have bought it, at a fraction of the cost and with none of the lead time. Put the options on one page: what each move does to the queue, what it costs, and when it lands. Smoothing usually wins on all three, and it keeps winning right up until the point where your variability is genuinely irreducible, at which stage adding capacity becomes the correct answer for the first time.

The objection you are already forming

"This is queueing theory borrowed from a factory. Our work is genuinely unpredictable. We do not control when a customer escalates, when a regulator asks a question, or when a build turns out to be harder than anybody thought. You cannot smooth demand you do not control, and telling me to release work on a cadence will not stop an incident happening on a Tuesday."

All true, and it concedes far less than it appears to. The question was never whether variability exists. It is what proportion of yours is imposed on you and what proportion you are manufacturing, and the honest audit is uncomfortable for nearly every portfolio that runs it.

Go back through the five sources. Your planning cadence is a choice. Whether unplanned work lands on the constraint or on an absorber is a choice. How many expedites you permit is a choice. How large you let a work item get before it enters the queue is a choice. How much you invest in reducing rework is a choice. Genuinely exogenous arrival variance is one contributor among several, and in most organizations it is not the largest. The escalation you cannot predict is real. The fact that it lands directly on your scarcest specialist, mid-task, with no triage layer in between, is a design decision, and one that was probably never consciously made at all.

And where the variability really is irreducible, the conclusion is not that the analysis fails. It is the opposite, and it is more urgent: if you cannot lower the variability term, then the same equation says you must lower the utilization term to compensate, or accept the queue and plan honestly around it. What you cannot do is run a high-variability system at 90 per cent utilization and treat the resulting queue as a performance problem to be solved by pressure. That system is behaving exactly as the arithmetic requires. Pushing harder on it only moves it further up the curve.

The deeper shift is in what you take a delay report to be telling you. A project that waited nine weeks at a gate has not necessarily met an under-resourced gate, a slow team, or a manager who needs to grip things. It has met a rhythm, produced by an approval calendar, an escalation policy, an expedite culture and a sizing convention, none of which appear anywhere on its plan and all of which are more changeable than headcount. Before you go and ask for more people, spend a fortnight working out how much of your queue is being generated by the way you release work into it. In most portfolios that is the largest number nobody has ever calculated, and it is the one you can act on this quarter.

This piece is the second half of The Utilization Trap: that one covers how full your constraint is, this one covers how evenly the work reaches it, and the two multiply. If you have not yet established which resource all of this applies to, start with How to Find the One Resource That Sets Your Delivery Speed. The cheapest single source of smoothing available to most portfolios is starting less work at once, and if you want to know what the queue you are currently tolerating is worth in money, it is drag cost that converts it.


Leave a Reply

Your email address will not be published. Required fields are marked *